Sammanfattning

This thesis evaluates the effectiveness of five multimodal large language models, GPT-4o, GPT-4.1, Gemini 2.0 Flash, Gemini 2.5 Flash, and Mistral Medium, for extracting structured information from Swedish housing association reports. The models are benchmarked on their ability to retrieve information from 19 predefined fields using two inference strategies: page-bypage and multi-page. The evaluation focuses on three dimensions: extraction performance, cost-efficiency, and strategy-related tradeoffs. The results reveal statistically significant differences in both cost and precision, with 95% confidence intervals confirming that the multi-page approach consistently outperforms page-by-page processing on both metrics. In contrast, differences in recall and F1 score were not statistically significant, with overlapping confidence intervals suggesting these metrics are more sensitive to model-specific behavior than strategy choice. GPT‑4.1 with multipage processing achieved the highest precision of 0.70 (95% CI: 0.56, 0.84), while Gemini 2.0 Flash offered the best cost-efficiency, being over 40 times more cost-effective than GPT-4.1. These findings support the viability of multimodal large language models for structured information extraction from domain-specific, image-based documents in non-English contexts. They also underscore the importance of choosing the right model-strategy combination based on performance, cost, and application requirements. The study concludes with a guide for real-world deployment of such systems in production environments.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.