paper-with-me

홈 › Papers

MUSCAT: MUltilingual, SCientific ConversATion Benchmark

2026-04-17 · Supriti Sinhamahapatra, Thai-Binh Nguyen, Yiğit Oğuz, Enes Ugan, Jan Niehues, Alexander Waibel arxiv

The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were a multilingual speaker. To create this experience, speech technology needs to address several challenges: Handling mixed multilingual input, specific vocabulary, and code-switching. However, there is currently no dataset benchmarking this situation. We propose a new benchmark to evaluate current Automatic Speech Recognition (ASR) systems, whether they are able to handle these challenges. The benchmark consists of bilingual discussions on scientific papers between multiple speakers, each conversing in a different language. We provide a standard evaluation framework, beyond Word Error Rate (WER) enabling consistent comparison of ASR performance across languages. Experimental results demonstrate that the proposed dataset is still an open challenge for state-of-the-art ASR systems. The dataset is available in https://huggingface.co/datasets/goodpiku/muscat-eval. Keywords: multilingual, speech recognition, audio segmentation, speaker diarization

📄 PDF Abstract BibTeX arXiv:2604.15929

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech Recognition

Similar Papers 제목 키워드 기반

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

2026-06-29 · M. K. Arabov arxiv

Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these communities, which affects international collaboration and the progres…

DISPLACE Challenge: DIarization of SPeaker and LAnguage in Conversational Environments

2023-03-01 · Shikha Baghel, Shreyas Ramoji, Sidharth, Ranjana H 외

In multilingual societies, social conversations often involve code-mixed speech. The current speech technology may not be well equipped to extract information from multi-lingual multi-speaker conversations. The DISPLACE …

speaker-diarizationSpeaker Diarization

ADIDA: Automatic Dialect Identification for Arabic

2019-06-01 · NAACL 2019 6 · Ossama Obeid, Mohammad Salameh, Houda Bouamor, Nizar Habash

This demo paper describes ADIDA, a web-based system for automatic dialect identification for Arabic text. The system distinguishes among the dialects of 25 Arab cities (from Rabat to Muscat) in addition to Modern Standar…

Dialect Identification

Code-switched inspired losses for spoken dialog representations

2021-11-01 · EMNLP 2021 11 · Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel

Spoken dialogue systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of code-switching). In this work, we introduce new pretraining losses tailored to learn gen…

RetrievalSpoken Dialogue Systems

Since the Scientific Literature Is Multilingual, Our Models Should Be Too

2024-03-27 · Abteen Ebrahimi, Kenneth Church

English has long been assumed the $\textit{lingua franca}$ of scientific research, and this notion is reflected in the natural language processing (NLP) research involving scientific document representation. In this posi…

DiversityPosition