paper-with-me

Papers

On Assessing the Relevance of Code Reviews Authored by Generative Models

2025-12-17 · Robert Heumüller, Frank Ortmeier arxiv

The use of large language models like ChatGPT in code review offers promising efficiency gains but also raises concerns about correctness and safety. Existing evaluation methods for code review generation either rely on automatic comparisons to a single ground truth, which fails to capture the variability of human perspectives, or on subjective assessments of "usefulness", a highly ambiguous concept. We propose a novel evaluation approach based on what we call multi-subjective ranking. Using a dataset of 280 self-contained code review requests and corresponding comments from CodeReview StackExchange, multiple human judges ranked the quality of ChatGPT-generated comments alongside the top human responses from the platform. Results show that ChatGPT's comments were ranked significantly better than human ones, even surpassing StackExchange's accepted answers. Going further, our proposed method motivates and enables more meaningful assessments of generative AI's performance in code review, while also raising awareness of potential risks of unchecked integration into review processes.

📄 PDF Abstract BibTeX arXiv:2512.15466

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Using Learning Progressions to Guide AI Feedback for Science Learning

2026-03-03 · Xin Xia, Nejla Yuruk, Yun Wang, Xiaoming Zhai arxiv

Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubrics authored by domain experts. While effective, rubric authoring is time…

The Impact of LLM-Generated Reviews on Recommender Systems: Textual Shifts, Performance Effects, and Strategic Platform Control

2025-11-02 · Itzhak Ziv, Moshe Unger, Hilah Geva arxiv

The rise of generative AI technologies is reshaping content-based recommender systems (RSes), which increasingly encounter AI-generated content alongside human-authored content. This study examines how the introduction o…

Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time

2026-04-01 · Razvan Mihai Popescu, David Gros, Andrei Botocan, Rahul Pandita 외 arxiv

The rise of large language models for code has reshaped software development. Autonomous coding agents, able to create branches, open pull requests, and perform code reviews, now actively contribute to real-world project…

A Literature Review of Literature Reviews in Pattern Analysis and Machine Intelligence

2024-02-20 · Penghai Zhao, Xin Zhang, Jiayue Cao, Ming-Ming Cheng 외

The rapid advancements in Pattern Analysis and Machine Intelligence (PAMI) have led to an overwhelming expansion of scientific knowledge, spawning numerous literature reviews aimed at collecting and synthesizing fragment…

ArticlesLanguage ModellingLarge Language ModelNavigate

Assessing Co-Authored Papers in Tenure Decisions: Implications for Research Independence and Career Strategies in Economics

2025-01-10 · Lekang Ren, Danyang Xie

In tenure decisions, the treatment of co-authored papers often raises questions about a candidate's research independence. This study examines the effects of solo versus collaborative authorship in high-profile Economics…