Uppsats
Deep Reinforcement Learning for Autonomous Drone Pursuit-Evasion : From Agile Quadrotor Chase to Sensor-Limited Pursuit in Cluttered Environments
Master-uppsats
Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
The proliferation of small, agile drones has increased the need for autonomous interceptors capable of pursuing evasive targets in cluttered environments, where onboard sensors are constrained by range and field of view. Classical guidance laws degrade when an adaptive opponent, physical obstacles, and partial observability appear simultaneously. Most reinforcement learning work on drone control addresses single-agent tasks rather than adversarial pursuit-evasion scenarios. This thesis investigates the extent to which deep reinforcement learning, combined with staged curricula and adversarial self-play, can produce effective pursuit-evasion policies for autonomous drones across a progression from an open arena with full-state observations to a cluttered, partially observable setting with an intelligent adversary. The investigation is structured around two studies in which policies are trained with Proximal Policy Optimisation (PPO) under an Asynchronous Multi-Stage Population-Based (AMSPB) self-play pipeline that combines scripted warm-up opponents, population-based opponent sampling, and alternating role training. Study I places two rigid-body quadrotors with full-state observations in an open arena. Study II introduces procedurally generated obstacles, limits the pursuer to a bearing-only camera and Lidar-like range sensors for obstacle detection, and trains the pursuer against a fully adversarial evader that exploits sensor blind spots and obstacle occlusion. In Study I, the pursuer captures its co-trained adversary in 99.7% of matches and generalises across all evaluated evader types. In Study II, the pursuer achieves capture in 82.3% of matches while observing the target for only 55% of each match, developing search, reacquisition, and tracking behaviour without explicit programming. The controlled progression between the studies indicates that partial observability and obstacle clutter are the dominant factors reducing policy performance. All experiments are conducted in simulation, and no sim-to-real transfer is attempted. Video demonstrations: Study I — https://www.youtube.com/watch?v=FKUcvFmgtT0 | Study II — https://www.youtube.com/watch?v=_SLH2pJH2Hg
Information
- Författare
- Fallaha, Asaad
- Lärosäte / institution
- Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
- Nyckelord
- ⌕Förstärkningsinlärning⌕Obstacle Avoidance⌕Proximal Policy Optimisation⌕einforcement Learning⌕Pursuit-Evasion⌕Autonomous Drones⌕Self-Play⌕Quadrotor Dynamics⌕Partial Observability⌕Multi-Agent Training⌕Förföljelse och undanflykt⌕Autonoma drönare⌕Självspel⌕Quadrotordynamik⌕Partiell observerbarhet⌕Multiagent träning⌕Hinderundvikande
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)
Persson, Hannes
Publicerad: 2024
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Morales Sundstedt, Michael
Publicerad: 2026
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Ge, Zilin
Publicerad: 2026
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Vincenova, Lenka
Publicerad: 2025
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Liljedahl, Carl
Publicerad: 2025
Master-uppsats, Linköpings universitet/Artificiell intelligens och integrerade datorsystem
Moberg, Oskar
Publicerad: 2025