paper-with-me

홈 › Papers

BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation

2026-01-26 · Peng Sun, Xiangyu Zhang, Duan Wu, Lu Tan, Jian Lin, He Yang, Qi Qian, Yikai Wang arxiv

Accurate evaluation of user satisfaction is critical for iterative development of conversational AI. However, for open-ended assistants, traditional A/B testing lacks reliable metrics: explicit feedback is sparse, while implicit metrics are ambiguous. To bridge this gap, we introduce BoRP (Bootstrapped Regression Probing), a scalable framework for high-fidelity satisfaction evaluation. Unlike generative approaches, BoRP leverages the geometric properties of LLM latent space. It employs a polarization-index-based bootstrapping mechanism to automate rubric generation and utilizes Partial Least Squares (PLS) to map hidden states to continuous scores. Experiments on industrial datasets show that BoRP (Qwen3-8B/14B) significantly outperforms generative baselines (even Qwen3-Max) in alignment with human judgments. Furthermore, BoRP reduces inference costs by orders of magnitude, enabling full-scale monitoring and highly sensitive A/B testing via CUPED.

📄 PDF Abstract BibTeX arXiv:2601.18253

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GaborPINN: Efficient physics informed neural networks using multiplicative filtered networks

2023-08-10 · Xinquan Huang, Tariq Alkhalifah

The computation of the seismic wavefield by solving the Helmholtz equation is crucial to many practical applications, e.g., full waveform inversion. Physics-informed neural networks (PINNs) provide functional wavefield s…

Bootstrapping Linear Models for Fast Online Adaptation in Human-Agent Collaboration

2024-04-16 · Benjamin A Newman, Chris Paxton, Kris Kitani, Henny Admoni

Agents that assist people need to have well-initialized policies that can adapt quickly to align with their partners' reward functions. Initializing policies to maximize performance with unknown partners can be achieved …

Human Agent CollaborationImitation Learningregression

Choice Between Partial Trajectories: Disentangling Goals from Beliefs

2024-10-30 · Henrik Marklund, Benjamin Van Roy

As AI agents generate increasingly sophisticated behaviors, manually encoding human preferences to guide these agents becomes more challenging. To address this, it has been suggested that agents instead learn preferences…

Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision

2026-02-23 · Kartik Kuckreja, Parul Gupta, Muhammad Haris Khan, Abhinav Dhall arxiv

Deepfake detection models often generate natural-language explanations, yet their reasoning is frequently ungrounded in visual evidence, limiting reliability. Existing evaluations measure classification accuracy but over…

DeepFake DetectionVisual Reasoning

Confident Neural Network Regression with Bootstrapped Deep Ensembles

2022-02-22 · Laurens Sluijterman, Eric Cator, Tom Heskes

With the rise of the popularity and usage of neural networks, trustworthy uncertainty estimation is becoming increasingly essential. One of the most prominent uncertainty estimation methods is Deep Ensembles (Lakshminara…

Prediction Intervalsregression