Uppsats

Deep Reinforcement Learning for Autonomous Drone Pursuit-Evasion : From Agile Quadrotor Chase to Sensor-Limited Pursuit in Cluttered Environments

Master-uppsats

Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)

Publicerad: 2026

Språk: Engelska

Sammanfattning

The proliferation of small, agile drones has increased the need for autonomous interceptors capable of pursuing evasive targets in cluttered environments, where onboard sensors are constrained by range and field of view. Classical guidance laws degrade when an adaptive opponent, physical obstacles, and partial observability appear simultaneously. Most reinforcement learning work on drone control addresses single-agent tasks rather than adversarial pursuit-evasion scenarios. This thesis investigates the extent to which deep reinforcement learning, combined with staged curricula and adversarial self-play, can produce effective pursuit-evasion policies for autonomous drones across a progression from an open arena with full-state observations to a cluttered, partially observable setting with an intelligent adversary. The investigation is structured around two studies in which policies are trained with Proximal Policy Optimisation (PPO) under an Asynchronous Multi-Stage Population-Based (AMSPB) self-play pipeline that combines scripted warm-up opponents, population-based opponent sampling, and alternating role training. Study I places two rigid-body quadrotors with full-state observations in an open arena. Study II introduces procedurally generated obstacles, limits the pursuer to a bearing-only camera and Lidar-like range sensors for obstacle detection, and trains the pursuer against a fully adversarial evader that exploits sensor blind spots and obstacle occlusion. In Study I, the pursuer captures its co-trained adversary in 99.7% of matches and generalises across all evaluated evader types. In Study II, the pursuer achieves capture in 82.3% of matches while observing the target for only 55% of each match, developing search, reacquisition, and tracking behaviour without explicit programming. The controlled progression between the studies indicates that partial observability and obstacle clutter are the dominant factors reducing policy performance. All experiments are conducted in simulation, and no sim-to-real transfer is attempted. Video demonstrations: Study I — https://www.youtube.com/watch?v=FKUcvFmgtT0 | Study II — https://www.youtube.com/watch?v=_SLH2pJH2Hg

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.