Uppsats

Integrating Vision–Language Models with Medical Imaging and Clinical Data for Improved Lymphoma Diagnosis

H

Chalmers tekniska högskola / Institutionen för elektroteknik

Publicerad: 2026

Språk: Engelska

Sammanfattning

Recent advancements in deep learning have fundamentally transformed the fieldof medical imaging. Along with the increasing burden of cancer worldwide, thedemand for automated diagnostic tools has surged. An example of such process isthe Lymphoma Artificial Reader System (LARS), an ensemble model composed of 10ResNet34-based models trained on over 17,000 [18F] FDG-PET/CT scans. AlthoughLARS has achieved impressive diagnostic accuracy, it relies solely on image featuresand lacks the capability to understand and utilize patient metadata (such as age,sex, and smoking history) and treatment information. However, these structuredclinical records play a crucial role in diagnostic decisions in daily clinical practice.In order to bridge this gap, we propose a new vision–language version of LARS(namely MM-LARS) capable of incorporating patient context via a tabular-totexttransformation. By converting structured patient data into natural languageprompts, we effectively fuse clinical text with dual-view PET images. Through extensiveablation experiments, we demonstrate that this multimodal approach notonly improves overall diagnostic performance over vision-only baselines, but alsosignificantly boosts model specificity. Notably, the inclusion of clinical text successfullycorrects false-positive misdiagnoses caused by ambiguous visual features. Ourwork provides practical insight into the next generation of multimodal diagnosticAI, offering a framework that is potentially generalizable beyond lymphoma.

Information

Författare
Tang, Yihan
Lärosäte / institution
Chalmers tekniska högskola / Institutionen för elektroteknik
Publiceringsdatum
2026
Uppsatstyp
H
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.