Uppsats
Formal Verification of Causal Reinforcement Learning Controllers
Master-uppsats
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publicerad: 2024
Språk: Engelska
Sammanfattning
Deep Reinforcement Learning (DRL) has been successfully applied to complex decision-making tasks in many fields, including safety-critical ones like autonomous driving, where safety and adherence to specifications are crucial. Nonetheless, standard DRL faces limitations, such as a lack of generalizability and inefficient training. Causal Reinforcement Learning (CRL) tackles these issues by learning robust structural knowledge from causal relations in the environment. Formal verification approaches can provide guarantees or identify violations of a DRL model’s behavior against safety properties. Both methods have individually demonstrated efficacy in improving DRL robustness and safety, but there is a research gap in connecting them. This thesis aims to address this research gap by studying the verifiability of CRL compared to DRL to provide insights into the potential of combining the approaches. We train a DRL and a CRL controller on an autonomous driving use case that optimizes efficient driving, stability, energy efficiency, and adherence to speed limits. The trained models are evaluated in various out-of-distribution settings. We find that the CRL controller requires only a fraction of the training episodes of the DRL controller and behaves more stably during training. In all experiments, the CRL controller generalizes better to new settings. We formally verify the trained models using model checking on formal models constructed from sampling data using abstractions. Examined properties cover the different aspects of the model’s performance. Both types of models can be verified using the same methodology. The formal model constructed from the CRL model is often smaller and requires less build time. The effect of CRL on stability and training efficiency is evident in early training phases, where the model checking results highlight the controller’s more stable behavior and more effective training. The DRL and CRL controllers that achieved the highest reward during training achieved similar results from model checking.
Information
- Författare
- Marie Schmidt, Jule
- Lärosäte / institution
- KTH/Skolan för elektroteknik och datavetenskap (EECS)
- Publiceringsdatum
- 2024
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Karlstads universitet/Institutionen för ingenjörsvetenskap och fysik (from 2013)
Persson, Hannes
Publicerad: 2024
Yrkesexamen på avancerad nivå, Luleå tekniska universitet/Institutionen för system- och rymdteknik
Andersson, Kevin
Publicerad: 2025
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Ali, Hamza
Publicerad: 2024
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Vincenova, Lenka
Publicerad: 2025
Kandidat-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Asghari, Yas, Arkhang, Dawa
Publicerad: 2025
Master-uppsats, Uppsala universitet/Institutionen för informationsteknologi
Nowosadko, Konrad
Publicerad: 2024