Uppsats

Evaluating Neural Network Architectures in Real-Time Deep Reinforcement Learning with Proximal Policy Optimisation in Counter-Strike 2

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2026

Språk: Engelska

Sammanfattning

Deep Reinforcement Learning (DRL) has achieved strong results in controlled research settings, yet its application to modern video games remains limited due to restricted interfaces, sparse rewards, and high computational demands. This thesis investigates the feasibility of training DRL agents in Counter- Strike 2, a first-person shooter characterised by partial observability, real-time decision-making, and limited access to game state information. The objective is to evaluate how neural network architecture influences the performance of Proximal Policy optimisation agents operating directly from visual input. Four architectures were examined: NatureCNN, NatureCNN + Long Short-Term Memory (LSTM), Importance Weighted Actor-Learner Architecture (IMPALA), and IMPALA + LSTM. A four stage curriculum learning strategy with progressively increased task difficulty was utilised. The training was conducted under constrained hardware conditions, which demonstrated feasibility with limited computational resources. Agents were evaluated in the deathmatch format using combat metrics, learning stability metrics, and behavioural analyses of navigational coverage and action distributions. Results reveal architectural trade-offs, where NatureCNNbased models showed broader exploration, greater action diversity, and more consistent convergence. IMPALA-based architectures achieved higher peak performance but exhibited reduced reliability and more concentrated policies. Recurrent variants provided a context-dependent advantage, with IMPALA + LSTM demonstrating improved combat effectiveness under higher environmental complexity. Despite these differences, all DRL agents performed below bot, human, and behaviourally cloned baselines, although the best performing models approached in-game bot performance in some metrics. Overall, architectural design affect both performance and learned behaviour in real-time, partially observable environments.

Information

Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.