Uppsats

Investigating Abstract Reasoning Successes and Failures of Large Reasoning Models on ARC-AGI

Master-uppsats

Uppsala universitet/Institutionen för lingvistik och filologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

This thesis examines the performance of four large reasoning models (LRMs) on a manually annotated subset of the Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) benchmark. We evaluate results across 6 reasoning categories and 3 complexity variables. Additionally, we compare 3 prompting conditions: a few-shot default baseline, and 2 dual-stage approaches where the model must first decompose the task into a stepwise algorithm then use it during inference. One of these approaches is zero shot (Algo-Test) and the other is few-shot (Train-Algo-Test). Quantitative results show that the models excel at problems involving simple mathematical rules and transformations (geometric/numerical) but struggle with multi-step compositional and relational/spatial tasks. Train-Algo-Test prompting generally increases accuracy, while Algo-Test lowers it because the zero-shot environment prevents verification of the algorithm’s correctness. The highest scoring models were both from the DeepSeek-V4 series, with V4-Flash even outperforming the larger V3.2 model. Traces generated for incorrect problems are significantly longer and have and lower lexical entropy, while rumination rate is a less consistent marker. A qualitative error analysis reveals that overfitting, task misidentification, incorrect rule implementation, and incomplete exploration are common error patterns in traces associated with failure. Our work provides insight into where and why models struggle with abstract reasoning and contributes to interpretability of LLM behavior.

Information

Författare
Garcia, Kai
Lärosäte / institution
Uppsala universitet/Institutionen för lingvistik och filologi
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.