paper-with-me

홈 › Papers

An Unsupervised Decontamination Procedure For Improving The Reliability Of Human Judgments

2011-12-01 · NeurIPS 2011 12 · Michael C. Mozer, Benjamin Link, Harold Pashler

Psychologists have long been struck by individuals' limitations in expressing their internal sensations, impressions, and evaluations via rating scales. Instead of using an absolute scale, individuals rely on reference points from recent experience. This _relativity of judgment_ limits the informativeness of responses on surveys, questionnaires, and evaluation forms. Fortunately, the cognitive processes that map stimuli to responses are not simply noisy, but rather are influenced by recent experience in a lawful manner. We explore techniques to remove sequential dependencies, and thereby _decontaminate_ a series of ratings to obtain more meaningful human judgments. In our formulation, the problem is to infer latent (subjective) impressions from a sequence of stimulus labels (e.g., movie names) and responses. We describe an unsupervised approach that simultaneously recovers the impressions and parameters of a contamination model that predicts how recent judgments affect the current response. We test our _iterated impression inference_, or I^3, algorithm in three domains: rating the gap between dots, the desirability of a movie based on an advertisement, and the morality of an action. We demonstrate significant objective improvements in the quality of the recovered impressions.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Informativeness

Similar Papers 제목 키워드 기반

Improving Human Judgments by Decontaminating Sequential Dependencies

2010-12-01 · NeurIPS 2010 12 · Michael C. Mozer, Harold Pashler, Matthew Wilder, Robert V. Lindsey 외

For over half a century, psychologists have been struck by how poor people are at expressing their internal sensations, impressions, and evaluations via rating scales. When individuals make judgments, they are incapable …

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

2023-11-08 · Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E. Gonzalez 외

Large language models are increasingly trained on all the data ever produced by humans. Many have raised concerns about the trustworthiness of public benchmarks due to potential contamination in pre-training or fine-tuni…

HumanEvalMMLU

When Benchmarks Leak: Inference-Time Decontamination for LLMs

2026-01-27 · Jianzhe Chai, Yu Zhe, Jun Sakuma arxiv

Benchmark-based evaluation is the de facto standard for comparing large language models (LLMs). However, its reliability is increasingly threatened by test set contamination, where test samples or their close variants le…

The Effect of Document Summarization on LLM-Based Relevance Judgments

2025-12-05 · Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro, Gianluca Demartini arxiv

Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Models (LLMs) have recently been proposed as …

Document SummarizationInformation RetrievalText Summarization

Predicting Human Similarity Judgments Using Large Language Models

2022-02-09 · Raja Marjieh, Ilia Sucholutsky, Theodore R. Sumers, Nori Jacoby 외

Similarity judgments provide a well-established method for accessing mental representations, with applications in psychology, neuroscience and machine learning. However, collecting similarity judgments can be prohibitive…