paper-with-me

Papers

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

2026-08-26 · Tim Schopf, Tobias Schreieder, Akiko Aizawa arxiv

Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent "medium novelty" bias.

📄 PDF Abstract BibTeX arXiv:2608.25660

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Small Language Models as Judges for Rubric-Based Reinforcement Learning

2026-08-30 · Fengyu Xie, Yilun Zhao, Bingsen Chen, Arman Cohan 외 hf

Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific criteria. However, this makes reward computation expensive: training r…

Reinforcement Learning

Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge

2026-06-12 · Shaojie Yin arxiv

Large language models (LLMs) are now widely used as automatic judges for open-ended instruction-following evaluation. This practice is convenient, scalable, and often more semantically aware than reference-based metrics,…

Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation

2025-12-23 · Bhaktipriya Radharapu, Eshika Saxena, Kenneth Li, Chenxi Whitehouse 외 arxiv

As LLM-based judges become integral to industry applications, obtaining well-calibrated uncertainty estimates efficiently has become critical for production deployment. However, existing techniques, such as verbalized co…

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

2026-06-08 · Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li 외 arxiv

While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-e…

Mathematical ReasoningStyle Transfer

What Do Llamas Really Think? Revealing Preference Biases in Language Model Representations

2023-11-30 · Raphael Tang, Xinyu Zhang, Jimmy Lin, Ferhan Ture

Do large language models (LLMs) exhibit sociodemographic biases, even when they decline to respond? To bypass their refusal to "speak," we study this research question by probing contextualized embeddings and exploring w…

Language ModelingLanguage Modelling