This paper introduces a set of techniques for learning game state evaluation functions through reinforcement learning. First, we generalize tree bootstrapping, i.e. learning the values of states encountered during search rather than restricting updates to states observed during matches, to the setting of reinforcement learning with non-linear function approximation. Second, we modifies Unbounded Best-First Minimax by extending best action sequences to terminal states. Third, we replace the traditional binary game outcome $+1/-1$ with richer reinforcement signals, including quick wins, delayed losses, and scoring. Fourth, we propose a completion mechanism that exploits state resolution.
Finally, we introduce a novel action-selection distribution, referred to as the ordinal distribution.
Experimental results show that each of these techniques contributes to substantial improvements in playing strength. We integrate them into a unified algorithm, Athénan, and compare it against ExIt, a leading self-play reinforcement learning approach without prior knowledge.
Our results demonstrate that Athénan consistently outperforms ExIt.
We further evaluate Athénan on the games Hex, Othello, and Arimaa, where it surpasses state-of-the-art performance without relying on domain-specific knowledge. In addition, we consider the single-player game Morpion Solitaire, in which Athénan again reaches state-of-the-art results under the same constraint.
Overall, these results show that reinforcement learning, when combined with the proposed techniques, can achieve state-of-the-art performance across a diverse range of games without the need for handcrafted heuristics or expert knowledge.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design | 0 | 7.34 | 17-08-2026 |
| 2 | Approximation-Free Differentiable Oblique Decision Trees | 0 | 9.6 | 17-08-2026 |
| 3 | End-to-End Deep Learning for Predicting Metric Space-Valued Outputs | 0 | 10.66 | 17-08-2026 |
| 4 | The Sample Complexity of Parameter-Free Stochastic Convex Optimization | 0 | 5.7 | 17-08-2026 |
| 5 | High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks | 0 | 8.7 | 17-08-2026 |
| 6 | Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective | 0 | 10.94 | 17-08-2026 |
| 7 | MO-Gymnasium - environments for reinforcement learning | 0 | 26.67 | 19-07-2026 |
| 8 | Near-optimal Delta-convex Estimation of Lipschitz Functions | 0 | 9.71 | 17-08-2026 |
| 9 | Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization | 0 | 7 | 17-08-2026 |
| 10 | Equitable Domination in Turiyam Graphs with Network Applications [version 1; peer review: 3 approved] | 0 | 7 | 01-06-2026 |