Kai Li, Qiang Fu, Hang Xu, Jian Cheng, Haobo Fu, Junliang Xing, Bingyun Liu
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfectinformation games. It decomposes the total regret into counterfactual regrets, utilizing local regret minimization algorithms, such as Regret Matching (RM) or RM+, to minimize them. Recent research establishes a connection between Online Mirror Descent (OMD) and RM+, paving the way for an optimistic variant PRM+ and its extension PCFR+. However, PCFR+ assigns uniform weights for each iteration when determining regrets, leading to substantial regrets when facing dominated actions. This work explores minimizing weighted counterfactual regret with optimistic OMD, resulting in a novel CFR variant PDCFR+. It integrates PCFR+ and Discounted CFR (DCFR) in a principled manner, swiftly mitigating negative effects of dominated actions and consistently leveraging predictions to accelerate convergence. Theoretical analyses prove that PDCFR+ converges to a Nash equilibrium, particularly under distinct weighting schemes for regrets and average strategies. Experimental results demonstrate PDCFR+'s fast convergence in common imperfect-information games. The code is available at https://github.com/ rpSebastian/PDCFRPlus.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2404.13891")
get_code_for_paper("2404.13891")
have("2404.13891")
Connect an agent — have() is free.