paper-with-me

홈 › Papers

When is it Better to Compare than to Score?

2014-06-25 · Nihar B. Shah, Sivaraman Balakrishnan, Joseph Bradley, Abhay Parekh, Kannan Ramchandran, Martin Wainwright

When eliciting judgements from humans for an unknown quantity, one often has the choice of making direct-scoring (cardinal) or comparative (ordinal) measurements. In this paper we study the relative merits of either choice, providing empirical and theoretical guidelines for the selection of a measurement scheme. We provide empirical evidence based on experiments on Amazon Mechanical Turk that in a variety of tasks, (pairwise-comparative) ordinal measurements have lower per sample noise and are typically faster to elicit than cardinal ones. Ordinal measurements however typically provide less information. We then consider the popular Thurstone and Bradley-Terry-Luce (BTL) models for ordinal measurements and characterize the minimax error rates for estimating the unknown quantity. We compare these minimax error rates to those under cardinal measurement models and quantify for what noise levels ordinal measurements are better. Finally, we revisit the data collected from our experiments and show that fitting these models confirms this prediction: for tasks where the noise in ordinal measurements is sufficiently low, the ordinal approach results in smaller errors in the estimation.

📄 PDF Abstract BibTeX arXiv:1406.6618

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Explicit Sampling Dependent Spectral Error Bound for Column Subset Selection

2015-05-04 · Tianbao Yang, Lijun Zhang, Rong Jin, Shenghuo Zhu

In this paper, we consider the problem of column subset selection. We present a novel analysis of the spectral norm reconstruction for a simple randomized algorithm and establish a new bound that depends explicitly on th…

BERTScore: Evaluating Text Generation with BERT

2019-04-21 · ICLR 2020 1 · Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 외

We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sen…

Image CaptioningMachine TranslationModel SelectionSentence+2

Trial-Based Dominance Enables Non-Parametric Tests to Compare both the Speed and Accuracy of Stochastic Optimizers

2022-12-19 · Kenneth V. Price, Abhishek Kumar, Ponnuthurai N Suganthan

Non-parametric tests can determine the better of two stochastic optimization algorithms when benchmarking results are ordinal, like the final fitness values of multiple trials. For many benchmarks, however, a trial can a…

BenchmarkingStochastic Optimization

Natural Language Generation enhances human decision-making with uncertain information

2016-06-10 · ACL 2016 8 · Dimitra Gkatzia, Oliver Lemon, Verena Rieser

Decision-making is often dependent on uncertain data, e.g. data associated with confidence scores or probabilities. We present a comparison of different information presentations for uncertain data and, for the first tim…

Decision MakingDecision Making Under UncertaintyText Generation

Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores

2025-09-27 · Ashwin Ramaswamy, Nestor Demeure, Ermal Rrapaj arxiv

New large language models (LLMs) are being released every day. Some perform significantly better or worse than expected given their parameter count. Therefore, there is a need for a method to independently evaluate model…