paper-with-me

Papers

Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation

2026-05-14 · GAng Peng arxiv

Holistic evaluation scores capture overall output quality but do not distinguish whether a model reproduced the structural form of a user's request from whether it preserved the user's specific intent. We propose a dimension-level intent fidelity evaluation framework, applied here through a structured prompt ablation study across 2,880 outputs spanning three languages, three task domains, and six LLMs, that separately measures structural recovery and intent fidelity for each semantic dimension. This framework reveals a systematic structural-fidelity split: among Chinese-language outputs with complete paired scores, 25.7% received perfect holistic alignment scores (GA=5) while exhibiting measurable dimensional intent deficits; among English-language outputs, this proportion rose to 58.6%. Human evaluation confirmed that these split-zone outputs represent genuine quality deficits and that dimensional fidelity scores track human judgements more reliably than holistic scores do. A public-private decomposition of 2,520 ablation cells characterises when models successfully compensate for missing intent and when they fail, while proxy annotation distinguishes prior inferability from default recoverability. A weight-perturbation experiment shows that moderate misalignment is typically absorbed, whereas severe dimensional inversion is consistently harmful. These findings demonstrate that dimension-level intent fidelity evaluation is a necessary complement to holistic assessment when evaluating LLM outputs for user-specific tasks.

📄 PDF Abstract BibTeX arXiv:2605.14517

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next

2026-03-12 · Ye Pan, Chi Kit Wong, Yuanhuiyi Lyu, Hanqian Li 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos rema…

How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity

2025-11-11 · Zihan Ma, Dongsheng Zhu, Shudong Liu, Taolin Zhang 외 arxiv

Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within complex tasks. We address this gap with a t…

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

2026-05-27 · Yongsik Seo, Wooseok Jeong, Eunyoung Kim, Hyeonseo Jang 외 arxiv

Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves. Millions of queries per day now pass through these systems, making c…

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions

2026-04-30 · Zhuoran Pan, Yue Li, Zhi Guan, Jianbin Hu 외 arxiv

The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of translating high-level user intents into functionally correct, state-dep…

FlexID: Training-Free Flexible Identity Injection via Intent-Aware Modulation for Text-to-Image Generation

2026-02-07 · Guandong Li, Yijun Ding arxiv

Personalized text-to-image generation aims to seamlessly integrate specific identities into textual descriptions. However, existing training-free methods often rely on rigid visual feature injection, creating a conflict …

Text-to-Image Generation