Uppsats
An Explainable AI and LLM-Assisted Framework for Within-Household SARS-CoV-2 Secondary Transmission Risk Prediction in Sweden
Master-uppsats
Lunds universitet/Institutionen för elektro- och informationsteknik
Publicerad: 2026
Språk: Engelska
Sammanfattning
Artificial intelligence (AI) is playing an expanding role in infectious disease surveillance and intervention planning by enabling analysis of large-scale health data. One of the key challenges in this area is to characterize transmission within households and to identify factors that influence both infectivity and susceptibility. The present study focuses on the COVID-19 pandemic, which not only generated unprecedented population-level data but also emphasized household secondary transmission as a central mechanism in the spread of respiratory infectious diseases. Predicting the household secondary transmission risk at both the household and individual levels can provide complementary evidence for targeted public health intervention. This thesis presents a multi-level explainable framework for predicting household SARS-CoV-2 secondary transmission risk, applied to the linked Swedish population and healthcare registers from the 2020 pre-vaccination period, covering 252,472 households and 608,473 individuals. The system integrates two independent prediction tasks within a unified framework: a household-level classifier that estimates the transmission probability, and an individual-level predictor that estimates relative susceptibility scores for household members. Four classifiers (logistic regression, random forest, XGBoost, and TabPFN) are benchmarked at both levels. TabPFN, a transformer-based foundation model pre-trained on causally structured synthetic datasets via in-context learning, and XGBoost show comparable performance at both the household level (AUC-ROC: 0.645 vs. 0.641) and the individual level (0.717 vs. 0.726). The focus of this study lies not in prediction per se, but in generating interpretable insights into the determinants of household transmission. Explainability is achieved through a structured three-level KernelSHAP analysis, i.e., global population-wide attribution, subgroup stratification, and local instance-level decomposition. Finally, an LLM-based explanation agent translates model predictions and SHAP attributions into plain-language narratives. Five frontier LLMs (Claude Sonnet 4.5, GPT-5.3, Grok 4.2, DeepSeek V3.2, and Llama 4) were evaluated using a two-tier framework combining qualitative and quantitative metrics, which demonstrated the real-world application potential of LLMs. The proposed framework integrates machine learning, explainable AI, and large language models to support more transparent and accessible public health decision-making. The code and materials for this work are publicly available at https://github.com/mingtongdu862-dot/covid19-household-transmission.
Information
- Författare
- Du, Mingtong
- Lärosäte / institution
- Lunds universitet/Institutionen för elektro- och informationsteknik
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, KTH/Skolan för teknikvetenskap (SCI)
Johanson, Filip
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Weidemann, Eivind Aksel
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Li, Hongyan
Publicerad: 2026
Master-uppsats, Lunds universitet/Matematik LTH
Brasar, Sparf Nils
Publicerad: 2026
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Saleh, Abdelrahman
Publicerad: 2026
Master-uppsats, Lunds universitet/Produktionsekonomi
Iveberg, Emma, Ekstrand, Erik
Publicerad: 2026