paper-with-me

Papers

LLMs can be easily Confused by Instructional Distractions

2025-02-05 · Yerin Hwang, Yongil Kim, Jahyun Koo, Taegwan Kang, Hyunkyung Bae, Kyomin Jung

Despite the fact that large language models (LLMs) show exceptional skill in instruction following tasks, this strength can turn into a vulnerability when the models are required to disregard certain instructions. Instruction-following tasks typically involve a clear task description and input text containing the target data to be processed. However, when the input itself resembles an instruction, confusion may arise, even if there is explicit prompting to distinguish between the task instruction and the input. We refer to this phenomenon as instructional distraction. In this paper, we introduce a novel benchmark, named DIM-Bench, specifically designed to assess LLMs' performance under instructional distraction. The benchmark categorizes real-world instances of instructional distraction and evaluates LLMs across four instruction tasks: rewriting, proofreading, translation, and style transfer -- alongside five input tasks: reasoning, code generation, mathematical reasoning, bias detection, and question answering. Our experimental results reveal that even the most advanced LLMs are susceptible to instructional distraction, often failing to accurately follow user intent in such cases.

📄 PDF Abstract BibTeX arXiv:2502.04362

Code (0)

등록된 구현이 없습니다.

Tasks

Bias DetectionCode GenerationInstruction FollowingMathematical ReasoningQuestion AnsweringStyle Transfer

Similar Papers 제목 키워드 기반

ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models

2026-06-25 · Xin Lin, Liang Zhang, Guoqi Ma, Hongyao Tu 외 arxiv

Open Relation Extraction (OpenRE) requires a model to extract unseen relations between head and tail entities from unstructured text for real-world applications. The core challenge of OpenRE lies in achieving reliable ge…

Relation Extraction

How Easily do Irrelevant Inputs Skew the Responses of Large Language Models?

2024-04-04 · Siye Wu, Jian Xie, Jiangjie Chen, Tinghui Zhu 외

By leveraging the retrieval of information from external knowledge databases, Large Language Models (LLMs) exhibit enhanced capabilities for accomplishing many knowledge-intensive tasks. However, due to the inherent flaw…

Retrieval

SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials

2026-04-28 · Zhaohui Li, Peng He, Zhiyuan Chen, Honglu Liu 외 arxiv

The need to evaluate instructional materials for K-12 science education has become increasingly important, as more educators use generative AI to create instructional materials. However, the review of instructional mater…

COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

2019-03-07 · CVPR 2019 6 · Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng 외

There are substantial instructional videos on the Internet, which enables us to acquire knowledge for completing various tasks. However, most existing datasets for instructional video analysis have the limitations in div…

Action Detection

Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents

2024-10-01 · Shiwei Wu, Chen Zhang, Yan Gao, Qimeng Wang 외

Instructional documents are rich sources of knowledge for completing various tasks, yet their unique challenges in conversational question answering (CQA) have not been thoroughly explored. Existing benchmarks have prima…

BenchmarkingConversational Question AnsweringQuestion Answering