Uppsats

Spatio-Temporal Attention in Online Visual Reinforcement Learning

Master-uppsats

Linnéuniversitetet/Institutionen för matematik och fysik (MF)

Publicerad: 2026

Språk: Engelska

Sammanfattning

This thesis investigates the integration of spatio-temporal attention mechanisms with distributional reinforcement learning, in the context of online visual reinforcement learning. While recent advances in reinforcement learning have enabled agents to learn directly from high-dimensional visual inputs, such approaches often suffer from poor sample efficiency, instability, and limited ability to capture long-range spatial and temporal dependencies. At the same time, Transformer-based architectures have demonstrated strong performance in sequential and visual modelling tasks, but their application in online reinforcement learning remains limited. To address this gap, this thesis introduces the Spatio-Temporal Quantile Network (STQN), a new architecture that integrates TimeSformer-based spatio-temporal attention with distributional reinforcement learning through Implicit Quantile Networks. The proposed model is designed to learn both rich visual representations and the distribution of future returns from pixel-based observations in an online setting. The approach is evaluated in the LunarLander-v3 environment using visual input, and compared against convolutional neural network (CNN) baselines, including Deep Q-Network (DQN), Categorical DQN (C51), and Implicit Quantile Networks (IQN), under a controlled experimental setup. The results indicate that the proposed Transformer-based architecture is capable of learning meaningful spatio-temporal representations and demonstrates competitive performance relative to CNN-based methods. In particular, the STQN agent shows potential improvements in learning stability and provides insights into how attention mechanisms influence representation learning in reinforcement learning systems. However, gains in sample efficiency and convergence speed are dependent on model complexity and training conditions, highlighting practical trade-off between representational power and computational cost. In addition, the thesis explores similarities between artificial agents and human gameplay behaviour by analysing learning progression and variability. The findings suggest that Transformer-based agents exhibit behavioural patterns that partially align with human strategies, although significant differences remain. Overall, this work contributes to a deeper understanding of how Transformer-based spatio-temporal modelling can be combined with distributional reinforcement learning in online environments. The results highlight both the potential and the limitations of attention-based approaches for improving learning efficiency, stability, and interpretability in visual reinforcement learning.

Information

Författare
Stenbom, Kristin
Lärosäte / institution
Linnéuniversitetet/Institutionen för matematik och fysik (MF)
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.