paper-with-me

Papers

Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework

2025-08-29 · Nils Dycke, Iryna Gurevych arxiv

Large Language Models (LLMs) have great potential to accelerate and support scholarly peer review and are increasingly used as fully automatic review generators (ARGs). However, potential biases and systematic errors may pose significant risks to scientific integrity; understanding the specific capabilities and limitations of state-of-the-art ARGs is essential. We focus on a core reviewing skill that underpins high-quality peer review: detecting faulty research logic. This involves evaluating the internal consistency between a paper's results, interpretations, and claims. We present a fully automated counterfactual evaluation framework that isolates and tests this skill under controlled conditions. Testing a range of ARG approaches, we find that, contrary to expectation, flaws in research logic have no significant effect on their output reviews. Based on our findings, we derive three actionable recommendations for future work and release our counterfactual dataset and evaluation framework publicly.

📄 PDF Abstract BibTeX arXiv:2508.21422

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework

2025-08-05 · Jialin Li, Jinzhe Li, Gengxu Li, Yi Chang 외 arxiv

With the advancement of code generation capabilities in large language models (LLMs), their reliance on input premises has intensified. When users provide inputs containing faulty premises, the probability of code genera…

Code Generation

Addressing Camera Sensors Faults in Vision-Based Navigation: Simulation and Dataset Development

2025-07-03 · Riccardo Gallon, Fabian Schiemenz, Alessandra Menicucci, Eberhard Gill arxiv

The increasing importance of Vision-Based Navigation (VBN) algorithms in space missions raises numerous challenges in ensuring their reliability and operational robustness. Sensor faults can lead to inaccurate outputs fr…

Minder: Faulty Machine Detection for Large-scale Distributed Model Training

2024-11-04 · Yangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang 외

Large-scale distributed model training requires simultaneous training on up to thousands of machines. Faulty machine detection is critical when an unexpected fault occurs in a machine. From our experience, a training tas…

Fault Detection

Zero-Shot Motor Health Monitoring by Blind Domain Transition

2022-12-12 · Serkan Kiranyaz, Ozer Can Devecioglu, Amir Alhams, Sadok Sassi 외

Continuous long-term monitoring of motor health is crucial for the early detection of abnormalities such as bearing faults (up to 51% of motor failures are attributed to bearing faults). Despite numerous methodologies pr…

Fault DetectionGenerative Adversarial Network

Optimal Detection of Faulty Traffic Sensors Used in Route Planning

2017-02-08 · Amin Ghafouri, Aron Laszka, Abhishek Dubey, Xenofon Koutsoukos

In a smart city, real-time traffic sensors may be deployed for various applications, such as route planning. Unfortunately, sensors are prone to failures, which result in erroneous traffic data. Erroneous data can advers…

Change Point DetectionGaussian Processes