Uppsats

Merging Language Reasoning into VLMs: A Study on Cross-Modal Capability Transfer

Master-uppsats

Uppsala universitet/Institutionen för lingvistik och filologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

Visual Language Models (VLMs) combine visual and textual inputs to generate text, often building upon the capabilities of Large Language Models(LLMs).While recent LLMs have acquired advanced reasoning abilities, such capabilities are not consistently present in current VLMs. This study explores whether reasoning behaviors can be transferred from Language Reasoning Models (LRMs) to VLMs through parameter-level model merging, without additional finetuning. We merge the reasoning-oriented LLM DeepSeek with the multimodal model LLaVA-NeXT and evaluate the resulting models on a range of multimodal reasoning benchmarks. Our experiments show that moderate fusion ratios can introduce reasoning-style behaviors such as structured generation and use of reasoning tokens and preserving core multimodal performance at the same time. In contrast, extreme fusion ratios tend to destabilize outputs and reduce accuracy. These findings suggest that parameter-level fusion offers a resource-efficientway for enhancing VLMs with reasoning capabilities. Key factors influencing success include architectural compatibility, tokenizer alignment, and prompting design, and future evaluation will refine our understanding of reasoning depth.

Information

Författare
Gao, Yan
Lärosäte / institution
Uppsala universitet/Institutionen för lingvistik och filologi
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska