paper-with-me

홈 › Papers

CUB: Benchmarking Context Utilisation Techniques for Language Models

2025-05-22 · Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein

Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) that encourage or suppress context utilisation have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) to help practitioners within retrieval-augmented generation (RAG) identify the best CMT for their needs. CUB allows for rigorous testing on three distinct context types, observed to capture key challenges in realistic context utilisation scenarios. With this benchmark, we evaluate seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to nine LMs. Our results show that most of the existing CMTs struggle to handle the full set of types of contexts that may be encountered in real-world retrieval-augmented scenarios. Moreover, we find that many CMTs display an inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples. Altogether, our results show the need for holistic tests of CMTs and the development of CMTs that can handle multiple context types.

📄 PDF Abstract BibTeX arXiv:2505.16518

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingFact CheckingQuestion AnsweringRAGRetrievalRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models

2025-10-03 · Jingyi Sun, Pepa Atanasova, Sagnik Ray Choudhury, Sekh Mainul Islam 외 arxiv

Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users, who cannot determine whether models draw…

A Reality Check on Context Utilisation for Retrieval-Augmented Generation

2024-12-22 · Lovisa Hagström, Sara Vera Marjanović, Haeun Yu, Arnav Arora 외

Retrieval-augmented generation (RAG) helps address the limitations of the parametric knowledge embedded within a language model (LM). However, investigations of how LMs utilise retrieved information of varying complexity…

Claim VerificationLanguage ModelingLanguage ModellingRAG+2

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

2026-06-17 · Sakshi Joshi, Dhruv Subhash Rathi, Sanskar Singh, Eldho Ittan George 외 arxiv

AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowle…

Speech Recognition

Le benchmarking de la reconnaissance d'entit\'es nomm\'ees pour le fran\ccais (Benchmarking for French NER)

2018-05-01 · JEPTALNRECITAL 2018 5 · Jungyeul Park

Cet article pr{\'e}sente une t{\^a}che du benchmarking de la reconnaissance de l{'}entit{\'e} nomm{\'e}e (REN) pour le fran{\c{c}}ais. Nous entrainons et {\'e}valuons plusieurs algorithmes d{'}{\'e}tiquetage de s{\'e}que…

BenchmarkingNER

Ph\oebus : un Logiciel d'Extraction de R\'eutilisations dans des Textes Litt\'eraires

2015-06-01 · JEPTALNRECITAL 2015 6 · Mohamed Amine Boukhaled, Zied Sellami, Jean-Gabriel Ganascia

Ph{\oe}bus est un logiciel d{'}extraction de r{\'e}utilisations dans des textes litt{\'e}raires. Il a {\'e}t{\'e} d{\'e}velopp{\'e} comme un outil d{'}analyse litt{\'e}raire assist{\'e}e par ordinateur. Dans ce contexte,…