paper-with-me

홈 › Papers

IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval

2025-03-06 · Tingyu Song, Guo Gan, Mingsheng Shang, Yilun Zhao

We introduce IFIR, the first comprehensive benchmark designed to evaluate instruction-following information retrieval (IR) in expert domains. IFIR includes 2,426 high-quality examples and covers eight subsets across four specialized domains: finance, law, healthcare, and science literature. Each subset addresses one or more domain-specific retrieval tasks, replicating real-world scenarios where customized instructions are critical. IFIR enables a detailed analysis of instruction-following retrieval capabilities by incorporating instructions at different levels of complexity. We also propose a novel LLM-based evaluation method to provide a more precise and reliable assessment of model performance in following instructions. Through extensive experiments on 15 frontier retrieval models, including those based on LLMs, our results reveal that current models face significant challenges in effectively following complex, domain-specific instructions. We further provide in-depth analyses to highlight these limitations, offering valuable insights to guide future advancements in retriever development.

📄 PDF Abstract BibTeX arXiv:2503.04644

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalInstruction FollowingRetrieval

Similar Papers 제목 키워드 기반

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

2023-10-31 · Yuxin Jiang, YuFei Wang, Xingshan Zeng, Wanjun Zhong 외

The ability to follow instructions is crucial for Large Language Models (LLMs) to handle various real-world applications. Existing benchmarks primarily focus on evaluating pure response quality, rather than assessing whe…

Instruction Following

IHEval: Evaluating Language Models on Following the Instruction Hierarchy

2025-02-12 · Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu 외

The instruction hierarchy, which establishes a priority order from system messages to user messages, conversation history, and tool outputs, is essential for ensuring consistent and safe behavior in language models (LMs)…

Instruction Following

IF-VidCap: Can Video Caption Models Follow Instructions?

2025-10-21 · Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei 외 arxiv

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, uncon…

Video CaptioningDense Captioning

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

2025-06-14 · Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos 외

Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or mul…

Beat TrackingGenre classificationInformation RetrievalInstruction Following+5

LIFEBench: Evaluating Length Instruction Following in Large Language Models

2025-05-22 · Wei zhang, Zhenhong Zhou, Junfeng Fang, Rongwu Xu 외

While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length instructions-e.g., write a 10,000-word nove…

Instruction FollowingText Generation