Uppsats
Evaluating Neural Network Architectures in Real-Time Deep Reinforcement Learning with Proximal Policy Optimisation in Counter-Strike 2
Master-uppsats
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publicerad: 2026
Språk: Engelska
Sammanfattning
Deep Reinforcement Learning (DRL) has achieved strong results in controlled research settings, yet its application to modern video games remains limited due to restricted interfaces, sparse rewards, and high computational demands. This thesis investigates the feasibility of training DRL agents in Counter- Strike 2, a first-person shooter characterised by partial observability, real-time decision-making, and limited access to game state information. The objective is to evaluate how neural network architecture influences the performance of Proximal Policy optimisation agents operating directly from visual input. Four architectures were examined: NatureCNN, NatureCNN + Long Short-Term Memory (LSTM), Importance Weighted Actor-Learner Architecture (IMPALA), and IMPALA + LSTM. A four stage curriculum learning strategy with progressively increased task difficulty was utilised. The training was conducted under constrained hardware conditions, which demonstrated feasibility with limited computational resources. Agents were evaluated in the deathmatch format using combat metrics, learning stability metrics, and behavioural analyses of navigational coverage and action distributions. Results reveal architectural trade-offs, where NatureCNNbased models showed broader exploration, greater action diversity, and more consistent convergence. IMPALA-based architectures achieved higher peak performance but exhibited reduced reliability and more concentrated policies. Recurrent variants provided a context-dependent advantage, with IMPALA + LSTM demonstrating improved combat effectiveness under higher environmental complexity. Despite these differences, all DRL agents performed below bot, human, and behaviourally cloned baselines, although the best performing models approached in-game bot performance in some metrics. Overall, architectural design affect both performance and learned behaviour in real-time, partially observable environments.
Information
- Författare
- Morales Sundstedt, Michael
- Lärosäte / institution
- KTH/Skolan för elektroteknik och datavetenskap (EECS)
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Uppsala universitet/Statistiska institutionen
Sundberg, Niklas
Publicerad: 2026
Master-uppsats, Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)
Fallaha, Asaad
Publicerad: 2026
Master-uppsats, Luleå tekniska universitet/Rymdteknik
Dyhr, Marcus Alexander
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Song, Dongfang
Publicerad: 2025
Master-uppsats, Linköpings universitet/Artificiell intelligens och integrerade datorsystem
Moberg, Oskar
Publicerad: 2025
Master-uppsats, Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)
Persson, Hannes
Publicerad: 2024