Uppsats

Wait, we can use AI for code reviews? : Investigating the usefulness of AI-generated code review comments

Kandidat-uppsats

Linköpings universitet/Institutionen för datavetenskap

Publicerad: 2025

Språk: Engelska

Sammanfattning

Recent advancements in artificial intelligence (AI), especially large language models (LLMs), have sparked significant interest in their potential applications. Code review is a widely adopted software development practice that helps participants understand the codebase and maintain software quality. However, this process is highly time-consuming, prompting interest in improving efficiency through advances in code review tools. This study investigates the applicability and effectiveness of LLMs as assistive tools in the code review process. We generate and systematically evaluate AI-generated code review comments, which are assessed by both the study authors (using scientifically derived criteria) and professional software engineers (reflecting industrial best practices). A review of scientific literature and a series of interviews with engineers reveal substantial theoretical alignment on what constitutes a good code review comment, particularly regarding the identification of functional errors, edge cases, maintainability concerns, and suggestions for more efficient solutions. However, analysis of inter-rater reliability and thematic categorization indicates a substantial divergence of what constitutes a useful code review comment between study authors and engineers in practice. Our results show that current state-of-the-art LLMs can generate code review comments that are often perceived as useful by both academics and practitioners. Specifically, 69.4% (95% CI [65.5%, 73.1%]) of comments were found useful by the study authors, and 92.8% (95% CI [88.9%, 95.4%]) by professional software engineers. While LLMs are not positioned to replace humans in the code review process, they show promise as effective assistive tools. We highlight challenges related to prompt engineering, subjectivity in usefulness assessments, and the potential of LLMs as a component in future code reviewing tools.

Information

Lärosäte / institution
Linköpings universitet/Institutionen för datavetenskap
Publiceringsdatum
2025
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.