paper-with-me

홈 › Papers

Speech Trax: A Bottom to the Top Approach for Speaker Tracking and Indexing in an Archiving Context

2016-05-01 · LREC 2016 5 · F{\'e}licien Vallet, Jim Uro, J{\'e}r{\'e}my Andriamakaoly, Hakim Nabi, Mathieu Derval, Jean Carrive

With the increasing amount of audiovisual and digital data deriving from televisual and radiophonic sources, professional archives such as INA, France{'}s national audiovisual institute, acknowledge a growing need for efficient indexing tools. In this paper, we describe the Speech Trax system that aims at analyzing the audio content of TV and radio documents. In particular, we focus on the speaker tracking task that is very valuable for indexing purposes. First, we detail the overall architecture of the system and show the results obtained on a large-scale experiment, the largest to our knowledge for this type of content (about 1,300 speakers). Then, we present the Speech Trax demonstrator that gathers the results of various automatic speech processing techniques on top of our speaker tracking system (speaker diarization, speech transcription, etc.). Finally, we provide insight on the obtained performances and suggest hints for future improvements.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Ctrax extensions for tracking in difficult lighting conditions

2014-09-25 · Ulrich Stern, Chung-Hui Yang

The fly tracking software Ctrax by Branson et al. is popular for positional tracking of animals both within and beyond the fly community. Ctrax was not designed to handle tracking in difficult lighting conditions with st…

On Barriers to Archival Audio Processing

2025-07-11 · Peter Sullivan, Muhammad Abdul-Mageed arxiv

In this study, we leverage a unique UNESCO collection of mid-20th century radio recordings to probe the robustness of modern off-the-shelf language identification (LID) and speaker recognition (SR) methods, especially wi…

Language IdentificationSpeaker Recognition

Unsupervised neural and Bayesian models for zero-resource speech processing

2017-01-03 · Herman Kamper

In settings where only unlabelled speech data is available, zero-resource speech technology needs to be developed without transcriptions, pronunciation dictionaries, or language modelling text. There are two central prob…

ClusteringLanguage ModellingRepresentation Learning

Jointly Tracking and Separating Speech Sources Using Multiple Features and the generalized labeled multi-Bernoulli Framework

2018-04-16

This paper proposes a novel joint multi-speaker tracking-and-separation method based on the generalized labeled multi-Bernoulli (GLMB) multi-target tracking filter, using sound mixtures recorded by microphones. Standard …

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

2025-08-28 · Muhammad Shakeel, Yui Sudo, Yifan Peng, Chyi-Jiunn Lin 외 arxiv

This paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (AS…

Speaker DiarizationSpeech RecognitionSpeech Separation