Uppsats

EVALUATING RELIABILITY IN LLM-BASED PROCEDURAL CONTENT GENERATION IN GAMES

Kandidat-uppsats

Umeå universitet/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Procedural content generation (PCG) is widely used in the gaming industry to automate level design, and Large Language Models (LLMs) have recently emerged as a flexible tool for generating structured content from natural language instructions. However, LLM-generated game content frequently suffers from generation errors, such as unplayable level layouts, which limits its practical reliability. This thesis evaluates the extent to which algorithmic post-processing validation can mitigate these errors compared to direct model generation. A large-scale comparative experiment was conducted using a 2D grid-based maze expansion task, comparing two task-specific fine-tuned 8B models running locally against a larger, general 72B model accessed via an inference API. Each model was evaluated over 8,400 iterations across various maze dimensions using playability, structural validity, and constraint adherence as primary metrics. The results show that introducing a single algorithmic validation retry based on A* pathfinding significantly increases playability across all models. The smaller, task-adapted 8B models consistently outperformed the non-fine-tuned 72B model in both direct generation and retry conditions. This indicates that while automated feedback retry loops are highly effective at correcting logical errors during inference, task-specific training can be more critical for maintaining complex structural constraints than raw parameter scale.

Information

Lärosäte / institution
Umeå universitet/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.