paper-with-me

홈 › Papers

CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells

2024-09-29 · Atharva Naik, Marcus Alenius, Daniel Fried, Carolyn Rose

The task of automated code review has recently gained a lot of attention from the machine learning community. However, current review comment evaluation metrics rely on comparisons with a human-written reference for a given code change (also called a diff). Furthermore, code review is a one-to-many problem, like generation and summarization, with many "valid reviews" for a diff. Thus, we develop CRScore - a reference-free metric to measure dimensions of review quality like conciseness, comprehensiveness, and relevance. We design CRScore to evaluate reviews in a way that is grounded in claims and potential issues detected in the code by LLMs and static analyzers. We demonstrate that CRScore can produce valid, fine-grained scores of review quality that have the greatest alignment with human judgment among open source metrics (0.54 Spearman correlation) and are more sensitive than reference-based metrics. We also release a corpus of 2.9k human-annotated review quality scores for machine-generated and GitHub review comments to support the development of automated metrics.

📄 PDF Abstract BibTeX arXiv:2409.19801

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Grounded AI for Code Review: Resource-Efficient Large-Model Serving in Enterprise Pipelines

2025-10-11 · Sayan Mandal, Hua Jiang arxiv

Automated code review adoption lags in compliance-heavy settings, where static analyzers produce high-volume, low-rationale outputs, and naive LLM use risks hallucination and incurring cost overhead. We present a product…

Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding

2026-02-25 · Wee Joe Tan, Zi Rui Lucas Lim, Shashank Durgad, Karim Obegi 외 arxiv

Evaluating web usability typically requires time-consuming user studies and expert reviews, which often limits iteration speed during product development, especially for small teams and agile workflows. We present Avenir…

Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance

2026-03-28 · Jyotsana Khatri, Manasi Patwardhan arxiv

Rebuttal generation is a critical component of the peer review process for scientific papers, enabling authors to clarify misunderstandings, correct factual inaccuracies, and guide reviewers toward a more accurate evalua…

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation

2026-04-19 · Rongsheng Hu, Runwei Guan, Yicheng Di, Jiayu Bao 외 arxiv

Manual annotation of high-quality visual question answering with grounding (VQA-G) datasets, which pair visual questions with evidential grounding, is crucial for advancing vision-language models (VLMs), but remains unsc…

Visual Question AnsweringVisual Grounding

GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians

2025-10-15 · Xiuyuan Chen, Tao Sun, Dexin Su, Ailing Yu 외 arxiv

Current benchmarks for AI clinician systems, often based on multiple-choice exams or manual rubrics, fail to capture the depth, robustness, and safety required for real-world clinical practice. To address this, we introd…