Scalable Population Training for Zero-Shot Coordination

2511.11083v1 cs.LG, cs.AI 2025-11-17

Авторы:

Bingyu Hui, Lebin Yu, Quanming Yao, Yunpeng Qu, Xudong Zhang, Jian Wang

Abstract

Zero-shot coordination(ZSC) has become a hot topic in reinforcement learning research recently. It focuses on the generalization ability of agents, requiring them to coordinate well with collaborators that are not seen before without any fine-tuning. Population-based training has been proven to provide good zero-shot coordination performance; nevertheless, existing methods are limited by computational resources, mainly focusing on optimizing diversity in small populations while neglecting the potential performance gains from scaling population size. To address this issue, this paper proposes the Scalable Population Training (ScaPT), an efficient training framework comprising two key components: a meta-agent that efficiently realizes a population by selectively sharing parameters across agents, and a mutual information regularizer that guarantees population diversity. To empirically validate the effectiveness of ScaPT, this paper evaluates it along with representational frameworks in Hanabi and confirms its superiority.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

Scalable Population Training for Zero-Shot Coordination

Авторы:

Abstract

Ссылки и действия

Связанные статьи

The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated ...

Interaction Tensor Shap

Generalization Beyond Benchmarks: Evaluating Learnable Protein-Ligand Scoring Fu...

IdealTSF: Can Non-Ideal Data Contribute to Enhancing the Performance of Time Ser...

How Ensemble Learning Balances Accuracy and Overfitting: A Bias-Variance Perspec...

Навигация