Uppsats
Spatio-Temporal Attention in Online Visual Reinforcement Learning
Master-uppsats
Linnéuniversitetet/Institutionen för matematik och fysik (MF)
Publicerad: 2026
Språk: Engelska
Sammanfattning
This thesis investigates the integration of spatio-temporal attention mechanisms with distributional reinforcement learning, in the context of online visual reinforcement learning. While recent advances in reinforcement learning have enabled agents to learn directly from high-dimensional visual inputs, such approaches often suffer from poor sample efficiency, instability, and limited ability to capture long-range spatial and temporal dependencies. At the same time, Transformer-based architectures have demonstrated strong performance in sequential and visual modelling tasks, but their application in online reinforcement learning remains limited. To address this gap, this thesis introduces the Spatio-Temporal Quantile Network (STQN), a new architecture that integrates TimeSformer-based spatio-temporal attention with distributional reinforcement learning through Implicit Quantile Networks. The proposed model is designed to learn both rich visual representations and the distribution of future returns from pixel-based observations in an online setting. The approach is evaluated in the LunarLander-v3 environment using visual input, and compared against convolutional neural network (CNN) baselines, including Deep Q-Network (DQN), Categorical DQN (C51), and Implicit Quantile Networks (IQN), under a controlled experimental setup. The results indicate that the proposed Transformer-based architecture is capable of learning meaningful spatio-temporal representations and demonstrates competitive performance relative to CNN-based methods. In particular, the STQN agent shows potential improvements in learning stability and provides insights into how attention mechanisms influence representation learning in reinforcement learning systems. However, gains in sample efficiency and convergence speed are dependent on model complexity and training conditions, highlighting practical trade-off between representational power and computational cost. In addition, the thesis explores similarities between artificial agents and human gameplay behaviour by analysing learning progression and variability. The findings suggest that Transformer-based agents exhibit behavioural patterns that partially align with human strategies, although significant differences remain. Overall, this work contributes to a deeper understanding of how Transformer-based spatio-temporal modelling can be combined with distributional reinforcement learning in online environments. The results highlight both the potential and the limitations of attention-based approaches for improving learning efficiency, stability, and interpretability in visual reinforcement learning.
Information
- Författare
- Stenbom, Kristin
- Lärosäte / institution
- Linnéuniversitetet/Institutionen för matematik och fysik (MF)
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)
Vrielink, Isabel
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för naturgeografi och ekosystemvetenskap
Yang, Jingfan
Publicerad: 2025
Master-uppsats, Uppsala universitet/Institutionen för informationsteknologi
Ahmed, Mohit Uddin
Publicerad: 2025
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Wang, Jianbo
Publicerad: 2025
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Vincenova, Lenka
Publicerad: 2025
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Mo, Fan, Lendrop, Oscar
Publicerad: 2025