paper-with-me

Papers

Adaptively profiling models with task elicitation

2025-03-03 · Davis Brown, Prithvi Balehannina, Helen Jin, Shreya Havaldar, Hamed Hassani, Eric Wong

Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a method that automatically builds new evaluations to profile model behavior. Task elicitation finds hundreds of natural-language tasks -- an order of magnitude more than prior work -- where frontier models exhibit systematic failures, in domains ranging from forecasting to online harassment. For example, we find that Sonnet 3.5 over-associates quantum computing and AGI and that o3-mini is prone to hallucination when fabrications are repeated in-context.

📄 PDF Abstract BibTeX arXiv:2503.01986

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationLanguage ModelingLanguage ModellingLegal Reasoning

Similar Papers 제목 키워드 기반

Adaptive Prompt Elicitation for Text-to-Image Generation

2026-02-04 · Xinyi Wen, Lena Hegemann, Xiaofu Jin, Shuai Ma 외 arxiv

Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Prompt Elicitation (APE), a technique that …

Text-to-Image Generation

Towards Personalized Conversational Sales Agents : Contextual User Profiling for Strategic Action

2025-03-28 · Tongyoung Kim, Jeongeun Lee, Soojin Yoon, Sunghwan Kim 외

Conversational Recommender Systems (CRSs) aim to engage users in dialogue to provide tailored recommendations. While traditional CRSs focus on eliciting preferences and retrieving items, real-world e-commerce interaction…

Decision MakingRecommendation Systems

Explainable Active Learning for Preference Elicitation

2023-09-01 · Furkan Cantürk, Reyhan Aydoğan

Gaining insights into the preferences of new users and subsequently personalizing recommendations necessitate managing user interactions intelligently, namely, posing pertinent questions to elicit valuable information ef…

Active LearningFood recommendation

RPS: Information Elicitation with Reinforcement Prompt Selection

2026-04-15 · Tao Wang, Jingyao Lu, Xibo Wang, Haonan Huang 외 arxiv

Large language models (LLMs) have shown remarkable capabilities in dialogue generation and reasoning, yet their effectiveness in eliciting user-known but concealed information in open-ended conversations remains limited.…

Reinforcement LearningDialogue Generation

Context-Aware Personality Inference in Dyadic Scenarios: Introducing the UDIVA Dataset

2020-12-28 · Cristina Palmero, Javier Selva, Sorina Smeureanu, Julio C. S. Jacques Junior 외

This paper introduces UDIVA, a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload. The da…