paper-with-me

홈 › Papers

Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback

2025-08-14 · Osama Mohammed Afzal, Preslav Nakov, Tom Hope, Iryna Gurevych arxiv

Novelty assessment is a central yet understudied aspect of peer review, particularly in high volume fields like NLP where reviewer capacity is increasingly strained. We present a structured approach for automated novelty evaluation that models expert reviewer behavior through three stages: content extraction from submissions, retrieval and synthesis of related work, and structured comparison for evidence based assessment. Our method is informed by a large scale analysis of human written novelty reviews and captures key patterns such as independent claim verification and contextual reasoning. Evaluated on 182 ICLR 2025 submissions with human annotated reviewer novelty assessments, the approach achieves 86.5% alignment with human reasoning and 75.3% agreement on novelty conclusions - substantially outperforming existing LLM based baselines. The method produces detailed, literature aware analyses and improves consistency over ad hoc reviewer judgments. These results highlight the potential for structured LLM assisted approaches to support more rigorous and transparent peer review without displacing human expertise. Data and code are made available.

📄 PDF Abstract BibTeX arXiv:2508.10795

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Position on LLM-Assisted Peer Review: Addressing Reviewer Gap through Mentoring and Feedback

2026-01-14 · JungMin Yun, JuneHyoung Kwon, MiHyeon Kim, YoungBin Kim arxiv

The rapid expansion of AI research has intensified the Reviewer Gap, threatening the peer-review sustainability and perpetuating a cycle of low-quality evaluations. This position paper critiques existing LLM approaches t…

Hadith computational science in the age of large language models: a critical narrative review

2026-06-18 · Md. Ashraful Haque, Riasat Islam hf

We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet p…

Benchmark for Research Theme Classification of Scholarly Documents

2022-10-01 · sdp (COLING) 2022 10 · Óscar E. Mendoza, Wojciech Kusa, Alaa El-Ebshihy, Ronin Wu 외

We present a new gold-standard dataset and a benchmark for the Research Theme Identification task, a sub-task of the Scholarly Knowledge Graph Generation shared task, at the 3rd Workshop on Scholarly Document Processing.…

ClassificationGraph Generation

CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?

2025-10-28 · Qing Zong, Jiayu Liu, Tianshi Zheng, Chunyang Li 외 arxiv

Accurate confidence calibration in Large Language Models (LLMs) is critical for safe use in high-stakes domains, where clear verbalized confidence enhances user trust. Traditional methods that mimic reference confidence …

Sentiment Analysis in Scholarly Book Reviews

2016-03-04 · Hussam Hamdan, Patrice Bellot, Frederic Bechet

So far different studies have tackled the sentiment analysis in several domains such as restaurant and movie reviews. But, this problem has not been studied in scholarly book reviews which is different in terms of review…

Sentiment Analysis