paper-with-me

Papers

Interactive Benchmarks

2026-03-05 · Baoqing Yue, Zihan Zhu, Yutong Han, Brian Fan, Qian Sun, Jichen Feng, Hufei Yang, Yifan Zhang, Mengdi Wang arxiv

Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evaluations rely on subjective judgments. We argue that a core aspect of intelligence is the ability to decide what information to acquire and how to use it effectively. We propose Interactive Benchmarks, a unified evaluation paradigm that assesses a model's reasoning ability through budgeted multi-turn interaction. We evaluate models under this framework in two settings: Interactive Proofs, where models interact with a judge to solve Logic, UI2Html, and Mathematics tasks under objective feedback; and Interactive Games, where models reason strategically to maximize long-horizon utilities. Our results show that interactive benchmarks provide a more robust assessment of this dimension of model intelligence, revealing substantial room for improvement in interactive scenarios.

📄 PDF Abstract BibTeX arXiv:2603.04737

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests

2025-02-20 · Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari 외

We examine three evaluation paradigms: large question-answering benchmarks (e.g., MMLU and BBH), interactive games (e.g., Signalling Games or Taboo), and cognitive tests (e.g., for working memory or theory of mind). Firs…

Logical ReasoningMMLUQuestion Answering

Interactive Evaluation Requires a Design Science

2026-05-18 · Keyang Xuan, Peiyang Song, Pan Lu, Pengrui Han 외 arxiv

AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other agents, while many evaluation practices …

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

Interactiveness Field in Human-Object Interactions

2022-04-16 · CVPR 2022 1 · Xinpeng Liu, Yong-Lu Li, Xiaoqian Wu, Yu-Wing Tai 외

Human-Object Interaction (HOI) detection plays a core role in activity understanding. Though recent two/one-stage methods have achieved impressive results, as an essential step, discovering interactive human-object pairs…

Human-Object Interaction DetectionObject

Interactive multiclass segmentation using superpixel classification

2015-10-12 · Bérengère Mathieu, Alain Crouzil, Jean-Baptiste Puel

This paper adresses the problem of interactive multiclass segmentation. We propose a fast and efficient new interactive segmentation method called Superpixel Classification-based Interactive Segmentation (SCIS). From a f…

ClassificationGeneral ClassificationInteractive SegmentationSegmentation