paper-with-me

Papers

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

2026-06-17 · Sakshi Joshi, Dhruv Subhash Rathi, Sanskar Singh, Eldho Ittan George, R J Hari, Kaushal Bhogale, Mitesh M. Khapra arxiv

AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.

📄 PDF Abstract BibTeX arXiv:2606.19157

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

CUB: Benchmarking Context Utilisation Techniques for Language Models

2025-05-22 · Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee 외

Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric…

BenchmarkingFact CheckingQuestion AnsweringRAG+2

Quel type de syst\`emes utiliser pour la transcription automatique du fran\ccais ? Les HMM font de la r\'esistance (What system for the automatic transcription of French in audiovisual broadcasts ?)

2020-06-01 · JEPTALNRECITAL 2020 6 · Paul Del{\'e}glise, Carole Lailler

Forts d{'}une utilisation couronn{\'e}e de succ{\`e}s en traduction automatique, les syst{\`e}mes end-to-end dont la sortie r{\'e}side en une suite de caract{\`e}res, ont vu leur utilisation {\'e}tendue {\`a} la transcri…

Ph\oebus : un Logiciel d'Extraction de R\'eutilisations dans des Textes Litt\'eraires

2015-06-01 · JEPTALNRECITAL 2015 6 · Mohamed Amine Boukhaled, Zied Sellami, Jean-Gabriel Ganascia

Ph{\oe}bus est un logiciel d{'}extraction de r{\'e}utilisations dans des textes litt{\'e}raires. Il a {\'e}t{\'e} d{\'e}velopp{\'e} comme un outil d{'}analyse litt{\'e}raire assist{\'e}e par ordinateur. Dans ce contexte,…

Voxtral

2025-07-17 · Alexander H. Liu, Andy Ehrenberg, Andy Lo, Clément Denoix 외

We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio b…

Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models

2025-10-03 · Jingyi Sun, Pepa Atanasova, Sagnik Ray Choudhury, Sekh Mainul Islam 외 arxiv

Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users, who cannot determine whether models draw…