Uppsats

Safe Multi-Robot Planning Via Long-Run Averages

H

Chalmers tekniska högskola / Institutionen för data och informationsteknik

Publicerad: 2026

Språk: Engelska

Sammanfattning

Constrained reinforcement learning in Markov decision processes (MDPs) has receivedincreasing attention for its use in sequential decision making problems with safetyrequirements. This study investigated safe planning via long run average rewardusing MDPs. This thesis uses grid-world environments and builds on the Triple-QAframework [1]. Three approaches are evaluated: A single-agent baseline and twomulti agent extensions, a trivial joint-state extension, and separate Q-table approach.The results show that the single agent algorithm reproduces the result found in theoriginal framework and serves as a reliable baseline. The joint state space extensionsuffers from poor scalability due to exponential growth in the state action space,and therefore does not achieve comparable reward per agent as the baseline. Incontrast, the separate Q-table approach scales significantly better and achieves a levelcomparable to the single agent case both in an environment with and without agentinteraction. Although the result of two agents with trivial extension and separateQ-table was achieved to satisfy the constraint, the test with three agents did notsatisfy for both algorithms.

Information

Lärosäte / institution
Chalmers tekniska högskola / Institutionen för data och informationsteknik
Publiceringsdatum
2026
Uppsatstyp
H
Språk
Engelska