IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech RecognitionSimilar Papers 제목 키워드 기반
CUB: Benchmarking Context Utilisation Techniques for Language Models
Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric…
BenchmarkingFact CheckingQuestion AnsweringRAG+2Quel type de syst\`emes utiliser pour la transcription automatique du fran\ccais ? Les HMM font de la r\'esistance (What system for the automatic transcription of French in audiovisual broadcasts ?)
Forts d{'}une utilisation couronn{\'e}e de succ{\`e}s en traduction automatique, les syst{\`e}mes end-to-end dont la sortie r{\'e}side en une suite de caract{\`e}res, ont vu leur utilisation {\'e}tendue {\`a} la transcri…
Ph\oebus : un Logiciel d'Extraction de R\'eutilisations dans des Textes Litt\'eraires
Ph{\oe}bus est un logiciel d{'}extraction de r{\'e}utilisations dans des textes litt{\'e}raires. Il a {\'e}t{\'e} d{\'e}velopp{\'e} comme un outil d{'}analyse litt{\'e}raire assist{\'e}e par ordinateur. Dans ce contexte,…
Voxtral
We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio b…
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users, who cannot determine whether models draw…