Uppsats

AI and Game Theory: Strategic decision making in Imperfect Information Environments : Evaluating AI-Driven Poker bots against common playstyles

Kandidat-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

This thesis is focused on the research area of artificial intelligence (AI) and game theory. Specifically, it explores strategic decision making in imperfect information environments through the development of an AI-driven poker agent for Heads-Up No-Limit Texas Hold’em (HUNL). The main problem explored in the thesis is how effectively algorithms based on Counterfactual Regret Minimization can improve the performance of Heads-Up Texas Hold’em poker bots by identifying strategic patterns, compared to rule-based or equity-based approaches. Additionally, we evaluate exploitation of fixed playstyles through an adjusted version of the Monte Carlo Counterfactual Regret Minimization algorithm. A challenge of the project was dealing with the massive number of possible game states, while only having access to modest hardware. The implementation of the Monte Carlo Counterfactual Regret Minimization agent was made using the OpenSpiel framework, which provided the poker environment and the foundation for the algorithm. To make the training computationally feasible, we implemented several types of abstractions. This is done to simplify the game state representation without losing essential strategic information. We evaluated the Monte Carlo Counterfactual Regret Minimization agent against a rule-based equity bot, various baseline bots and Slumbot, a strong open-source poker agent with close to optimal strategy. The evaluation showed that the agent was significantly stronger than both the equity-based bot and the baseline bots. Our tests also showed how the agent continuously improved its strategy and gained stronger results when increasing the number of iterations during training. Although both our agent and the equity bot lost to Slumbot, our approach achieved a higher win rate and smaller chip losses. This indicates that algorithms based on Counterfactual Regret Minimization, even with some computational limitations, are a better strategy for solving poker and other imperfect information environments.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.