Uppsats

Improving representation learning in MARL using teammate-focused auxiliary task

Master-uppsats

Mälardalens universitet

Publicerad: 2026

Språk: Engelska

Sammanfattning

In cooperative Multi-Agent Reinforcement Learning (MARL), agents jointly interact with theenvironment and receive a shared team reward. However, learning cooperative behaviour that im-proves both performance and generalisation remains a difficult problem, especially when rewardsare sparse and execution is decentralised. Although value-based methods such as QMIX addressparts of this challenge, effective team play and generalisation remain challenging. This thesis ex-plores team play and generalisation in a cooperative MARL setting and introduces Teammate Ac-tion Forecast-QMIX (TAF-QMIX), a QMIX-based architecture augmented with Teammate ActionForecast (TAF) as an auxiliary task. TAF takes the recurrent agents’ hidden states as input andpredicts the next greedy action selected by each teammate’s current policy. The aim is to increasesample efficiency and encourage the agents’ hidden representations to capture teammate-relevantinformation, thereby improving team play while reducing reliance on opponent-specific modellingthat may limit generalisation. The architecture was trained and evaluated in the Half Field Of-fence (HFO) robot-soccer environment across several scenarios. The results indicated more stableattention to teammate-related and environment-related information, as well as a higher ratio ofsuccessful passes relative to pass attempts. However, the results did not indicate any consistentimprovement in training or evaluation performance. This thesis therefore concludes that the be-havioural differences encouraged by TAF did not, in this case, translate into reliable gains in goalrate or generalisation to unseen opponents.

Information

Lärosäte / institution
Mälardalens universitet
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.