Deep (Predictive) Discounted Counterfactual Regret Minimization

2511.08174v1 cs.LG, cs.AI, cs.GT 2025-11-15
Авторы:

Hang Xu, Kai Li, Haobo Fu, Qiang Fu, Junliang Xing, Jian Cheng

Abstract

Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers use neural networks to approximate its behavior. However, existing methods are mainly based on vanilla CFR and struggle to effectively integrate more advanced CFR variants. In this work, we propose an efficient model-free neural CFR algorithm, overcoming the limitations of existing methods in approximating advanced CFR variants. At each iteration, it collects variance-reduced sampled advantages based on a value network, fits cumulative advantages by bootstrapping, and applies discounting and clipping operations to simulate the update mechanisms of advanced CFR variants. Experimental results show that, compared with model-free neural algorithms, it exhibits faster convergence in typical imperfect-information games and demonstrates stronger adversarial performance in a large poker game.

Ссылки и действия

Связанные статьи

SpinGPT: A Large-Language-Model Approach to Playing Poker Correctly

## Контекст Область исследования — искусственный интеллект (ИИ) в играх, специально в покере. Игры, которым необходима с...

2025-09-30

From Leiden to Pleasure Island: The Constant Potts Model for Community Detection...

#### Контекст Community detection является одной из основных задач в области data science, состоящей в разбиении узлов г...

2025-09-06

Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Conte...

## Контекст Инверсное обучение наград (IRL) в играх с многими агентами (mean field games, MFGs) является важной задачей ...

2025-09-06