Uppsats

Formal Verification of Causal Reinforcement Learning Controllers

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2024

Språk: Engelska

Sammanfattning

Deep Reinforcement Learning (DRL) has been successfully applied to complex decision-making tasks in many fields, including safety-critical ones like autonomous driving, where safety and adherence to specifications are crucial. Nonetheless, standard DRL faces limitations, such as a lack of generalizability and inefficient training. Causal Reinforcement Learning (CRL) tackles these issues by learning robust structural knowledge from causal relations in the environment. Formal verification approaches can provide guarantees or identify violations of a DRL model’s behavior against safety properties. Both methods have individually demonstrated efficacy in improving DRL robustness and safety, but there is a research gap in connecting them. This thesis aims to address this research gap by studying the verifiability of CRL compared to DRL to provide insights into the potential of combining the approaches. We train a DRL and a CRL controller on an autonomous driving use case that optimizes efficient driving, stability, energy efficiency, and adherence to speed limits. The trained models are evaluated in various out-of-distribution settings. We find that the CRL controller requires only a fraction of the training episodes of the DRL controller and behaves more stably during training. In all experiments, the CRL controller generalizes better to new settings. We formally verify the trained models using model checking on formal models constructed from sampling data using abstractions. Examined properties cover the different aspects of the model’s performance. Both types of models can be verified using the same methodology. The formal model constructed from the CRL model is often smaller and requires less build time. The effect of CRL on stability and training efficiency is evident in early training phases, where the model checking results highlight the controller’s more stable behavior and more effective training. The DRL and CRL controllers that achieved the highest reward during training achieved similar results from model checking.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.