Uppsats
Safe Multi-Robot Planning Via Long-Run Averages
H
Chalmers tekniska högskola / Institutionen för data och informationsteknik
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Constrained reinforcement learning in Markov decision processes (MDPs) has receivedincreasing attention for its use in sequential decision making problems with safetyrequirements. This study investigated safe planning via long run average rewardusing MDPs. This thesis uses grid-world environments and builds on the Triple-QAframework [1]. Three approaches are evaluated: A single-agent baseline and twomulti agent extensions, a trivial joint-state extension, and separate Q-table approach.The results show that the single agent algorithm reproduces the result found in theoriginal framework and serves as a reliable baseline. The joint state space extensionsuffers from poor scalability due to exponential growth in the state action space,and therefore does not achieve comparable reward per agent as the baseline. Incontrast, the separate Q-table approach scales significantly better and achieves a levelcomparable to the single agent case both in an environment with and without agentinteraction. Although the result of two agents with trivial extension and separateQ-table was achieved to satisfy the constraint, the test with three agents did notsatisfy for both algorithms.
Information
- Författare
- Embaye, Eyob, Daun, Johan
- Lärosäte / institution
- Chalmers tekniska högskola / Institutionen för data och informationsteknik
- Publiceringsdatum
- 2026
- Uppsatstyp
- H
- Språk
- Engelska