Uppsats

Reinforcement Learning for Large Language Models

Master-uppsats

Uppsala universitet/Institutionen för informationsteknologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

Reinforcement learning (RL) has become a promising method for enhancing the reasoning capabilities of large language models (LLMs). However, significant questions persist regarding the extent to which RL-induced skills generalize beyond the training distribution. This thesis examines generalization across three dimensions: task variation, prompt variation, and environment variation, utilizing Group Relative Policy Optimization (GRPO) applied to Qwen3-4B-Thinking in mathematical reasoning tasks. Three training conditions are evaluated: GRPO with standard reward (GRPO-MATH), GRPO with tool-integrated reasoning (GRPO-MATH-TIR), and GRPO with tool-integrated reasoning combined with process rewards (GRPO-MATH-TIR+PR). All models are trained on the MATH dataset and evaluated on held-out MATH problems, paraphrased and Bengali-translated variants, AIME 2026, the Game of 24 combinatorial reasoning task, and ZebraLogic constraint satisfaction puzzles. Training is performed on the Leonardo HPC cluster (EuroHPC JU) using a fully sharded, data-parallel configuration with the VERL and verl-tool frameworks.

Information

Författare
Akhter, Samiha
Lärosäte / institution
Uppsala universitet/Institutionen för informationsteknologi
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.