paper-with-me

홈 › Papers

BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)

2025-10-14 · Tomas Ruiz, Siyao Peng, Barbara Plank, Carsten Schwemmer arxiv

Test-time scaling is a family of techniques to improve LLM outputs at inference time by performing extra computation. To the best of our knowledge, test-time scaling has been limited to domains with verifiably correct answers, like mathematics and coding. We transfer test-time scaling to the LeWiDi-2025 tasks to evaluate annotation disagreements. We experiment with three test-time scaling methods: two benchmark algorithms (Model Averaging and Majority Voting), and a Best-of-N sampling method. The two benchmark methods improve LLM performance consistently on the LeWiDi tasks, but the Best-of-N method does not. Our experiments suggest that the Best-of-N method does not currently transfer from mathematics to LeWiDi tasks, and we analyze potential reasons for this gap.

📄 PDF Abstract BibTeX arXiv:2510.12516

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning

2025-09-11 · Daniil Ignatev, Nan Li, Hugh Mee Wong, Anh Dang 외 arxiv

This system paper presents the DeMeVa team's approaches to the third edition of the Learning with Disagreements shared task (LeWiDi 2025; Leonardelli et al., 2025). We explore two directions: in-context learning (ICL) wi…

LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task

2025-10-09 · Elisa Leonardelli, Silvia Casola, Siyao Peng, Giulia Rizzi 외 arxiv

Many researchers have reached the conclusion that AI models should be trained to be aware of the possibility of variation and disagreement in human judgments, and evaluated as per their ability to recognize such variatio…

Natural Language InferenceParaphrase IdentificationSarcasm Detection

SemEval-2023 Task 11: Learning With Disagreements (LeWiDi)

2023-04-28 · Elisa Leonardelli, Alexandra Uma, Gavin Abercrombie, Dina Almanea 외

NLP datasets annotated with human judgments are rife with disagreements between the judges. This is especially true for tasks depending on subjective judgments such as sentiment analysis or offensive language detection. …

Sentiment Analysis

Which measure for PFE? The Risk Appetite Measure, A

2015-12-19

Potential Future Exposure (PFE) is a standard risk metric for managing business unit counterparty credit risk but there is debate on how it should be calculated. The debate has been whether to use one of many historical …

Optimal Strategies for the Decumulation of Retirement Savings under Differing Appetites for Liquidity and Investment Risks

2023-12-22 · Benjamin Avanzi, Lewis de Felice

A retiree's appetite for risk is a common input into the lifetime utility models that are traditionally used to find optimal strategies for the decumulation of retirement savings. In this work, we consider a retiree with…