paper-with-me

Papers

AG-LSEC: Audio Grounded Lexical Speaker Error Correction

2024-06-25 · Rohit Paturi, Xiang Li, Sundararajan Srinivasan

Speaker Diarization (SD) systems are typically audio-based and operate independently of the ASR system in traditional speech transcription pipelines and can have speaker errors due to SD and/or ASR reconciliation, especially around speaker turns and regions of speech overlap. To reduce these errors, a Lexical Speaker Error Correction (LSEC), in which an external language model provides lexical information to correct the speaker errors, was recently proposed. Though the approach achieves good Word Diarization error rate (WDER) improvements, it does not use any additional acoustic information and is prone to miscorrections. In this paper, we propose to enhance and acoustically ground the LSEC system with speaker scores directly derived from the existing SD pipeline. This approach achieves significant relative WDER reductions in the range of 25-40% over the audio-based SD, ASR system and beats the LSEC system by 15-25% relative on RT03-CTS, Callhome American English and Fisher datasets.

📄 PDF Abstract BibTeX arXiv:2406.17266

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingspeaker-diarizationSpeaker Diarization

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction

2023-06-15 · Rohit Paturi, Sundararajan Srinivasan, Xiang Li

Speaker diarization (SD) is typically used with an automatic speech recognition (ASR) system to ascribe speaker labels to recognized words. The conventional approach reconciles outputs from independently optimized ASR an…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models

2025-01-14 · Anurag Kumar, Rohit Paturi, Amber Afshan, Sundararajan Srinivasan

Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly d…

speaker-diarizationSpeaker Diarization

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

2023-09-11 · Tae Jin Park, Kunal Dhawan, Nithin Koluguri, Jagadeesh Balam

Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization that incorporates the prowess of LLMs to…

speaker-diarizationSpeaker Diarization

AV-Dialog: Spoken Dialogue Models with Audio-Visual Input

2025-11-14 · Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota arxiv

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues …

Boundary Detection

Detection of Lexical Stress Errors in Non-Native (L2) English with Data Augmentation and Attention

2020-12-29 · Daniel Korzekwa, Roberto Barra-Chicote, Szymon Zaporowski, Grzegorz Beringer 외

This paper describes two novel complementary techniques that improve the detection of lexical stress errors in non-native (L2) English speech: attention-based feature extraction and data augmentation based on Neural Text…

Data Augmentationtext-to-speechText to Speech