Uppsats
Evaluation of Data-Driven Report Generation Pipelines : Large Language Models for Decision Support in Smart Agriculture
Master-uppsats
Stockholms universitet/Institutionen för lingvistik
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
This thesis investigates the impact of different “semantic structured text to Large Language Model (LLM)” pipelines on the quality of generated out- puts within the context of greenhouse decision support. Consequently, this study addresses the existing research gap regarding agricultural LLMs and provides a Minimum Viable Product (MVP) level implementation for auto- mated report generation. For report generation, the research evaluates two methods of extracting semantic structured text: anomaly detection and statistical summaries, and compares the efficacy of two types of LLM: domain-specific LLM (Sinong) and general-purpose LLM (Gemini). To evaluate these configurations, each pipeline was utilized to generate ten daily reports for a selected greenhouse. The objective of these reports was to summarise the status of the farm and what needs attention. Sub- sequently, an LLM-as-a-Judge framework (powered by Claude) was em- ployed to evaluate the generated reports across four dimensions: Reason- ing quality, Scientific Accuracy, Actionability, and Usefulness. The results show that the pipeline integrating statistical summaries with the general-purpose LLM (Gemini) achieved the highest overall evaluation scores. Notably, the overall quality of the generated reports is primarily determined by the choice of the LLM, whereas the method of structuring the semantic input (anomaly vs. statistical) does not yield a statistically significant difference. Furthermore, a case study of this optimal pipeline was conducted to examine how the availability of ground truth influences the LLM-as-a-Judge evaluation mechanism. The findings reveal that providing ground truth yields more consistent scoring over most dimensions; the presence of ground truth fundamentally shifts the certain evaluation criteria of LLM-as Judge. Besides, a qualitative human evaluation validated these findings, demonstrating that while domain- specific pipelines exhibited ’limited vision’ in the report, the optimal Gemini pipeline maintained factual completeness across the entire report cycle.
Information
- Författare
- Zhou, Jun
- Lärosäte / institution
- Stockholms universitet/Institutionen för lingvistik
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Göteborgs universitet/Graduate School
Törnqvist, Cajsa, Rydh, Samuel
Publicerad: 2026-07-09
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Wu, Bingjie
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Hu, Junqi
Publicerad: 2026
Master-uppsats, Blekinge Tekniska Högskola/Institutionen för datavetenskap
Ahmed, Gazi Samia
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Nallapati, Sriram Kumar
Publicerad: 2026
Master-uppsats, Högskolan i Skövde/Institutionen för informationsteknologi
Nazir, Muhammad Sajid
Publicerad: 2026