paper-with-me

홈 › Papers

StakeBench: Evaluating Language Understanding Grounded in Market Commitment

2026-05-25 · Yunhua Pei, Jingyu Hu, Yiwei Shi, Hongnan Ma, Weiru Liu, John Cartlidge arxiv

Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers have committed to in the market. We introduce StakeBench, an evaluation framework for language understanding grounded in market commitment. StakeBench links 560,876 comments from 2,261 resolved markets to verified position, action, and market-odds records across Polymarket and Manifold. Supervision is derived from observable market behavior. Position sides, post-comment trading actions, and market-odds trajectories replace human annotation. Four diagnostic tasks test whether models detect market commitment, identify the revealed side, anticipate future action, and perform collective odds projection. Three commitment-aware metrics measure alignment with revealed preferences rather than perceived sentiment. Validity audits and explicit interpretation boundaries help distinguish observable commitment signals from latent belief and causal market-odds impact. Across 15 LLMs and 18 topics and platform settings, models partially recover position-side signals, with Directed Accuracy from 0.506 to 0.599, but show structural failures on later tasks. Ten of the fifteen models collapse to one or two action labels in future action anticipation, and no model consistently improves on the naive odds-direction baseline in collective odds projection. Model scale is not correlated with performance, finance-domain tuning does not improve revealed-side identification, and platform incentives strongly shape higher-order results. StakeBench is packaged with evaluation code and dataset under CC-BY 4.0.

📄 PDF Abstract BibTeX arXiv:2605.26074

Code (0)

등록된 구현이 없습니다.

Tasks

Action Anticipation

Similar Papers 제목 키워드 기반

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

2026-03-22 · Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Rubaya Tabassum 외 arxiv

Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVers…

Visual Question AnsweringVisual Grounding

ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models

2024-06-06 · Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang 외

Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. This development underscores the urgent need for evaluating value orientations and understanding of LLMs to …

Challenges and Prospects in Vision and Language Research

2019-04-19 · Kushal Kafle, Robik Shrestha, Christopher Kanan

Language grounded image understanding tasks have often been proposed as a method for evaluating progress in artificial intelligence. Ideally, these tasks should test a plethora of capabilities that integrate computer vis…

Natural Language Understanding

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

2025-02-04 · Xueqing Deng, Qihang Yu, Ali Athar, Chenglin Yang 외

This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome…

Image CaptioningPanoptic SegmentationSegmentation

Kestrel: Point Grounding Multimodal LLM for Part-Aware 3D Vision-Language Understanding

2024-05-29 · Junjie Fei, Mahmoud Ahmed, Jian Ding, Eslam Mohamed BAKR 외

While 3D MLLMs have achieved significant progress, they are restricted to object and scene understanding and struggle to understand 3D spatial structures at the part level. In this paper, we introduce Kestrel, representi…

Scene UnderstandingSegmentation