paper-with-me

홈 › Papers

LLM-based speaker diarization correction: A generalizable approach

2024-06-07 · Georgios Efstathiadis, Vijay Yadav, Anzar Abbas

Speaker diarization is necessary for interpreting conversations transcribed using automated speech recognition (ASR) tools. Despite significant developments in diarization methods, diarization accuracy remains an issue. Here, we investigate the use of large language models (LLMs) for diarization correction as a post-processing step. LLMs were fine-tuned using the Fisher corpus, a large dataset of transcribed conversations. The ability of the models to improve diarization accuracy in a holdout dataset from the Fisher corpus as well as an independent dataset was measured. We report that fine-tuned LLMs can markedly improve diarization accuracy. However, model performance is constrained to transcripts produced using the same ASR tool as the transcripts used for fine-tuning, limiting generalizability. To address this constraint, an ensemble model was developed by combining weights from three separate models, each fine-tuned using transcripts from a different ASR tool. The ensemble model demonstrated better overall performance than each of the ASR-specific models, suggesting that a generalizable and ASR-agnostic approach may be achievable. We have made the weights of these models publicly available on HuggingFace at https://huggingface.co/bklynhlth.

📄 PDF Abstract BibTeX arXiv:2406.04927

Code (1)

GeorgeEfstathiadis/LLM-Diarize-ASR-Agnostic 공식 구현

Tasks

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

DiaCorrect: End-to-end error correction for speaker diarization

2022-10-31 · Jiangyu Han, Yuhang Cao, Heng Lu, Yanhua Long

In recent years, speaker diarization has attracted widespread attention. To achieve better performance, some studies propose to diarize speech in multiple stages. Although these methods might bring additional benefits, m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction

2023-06-15 · Rohit Paturi, Sundararajan Srinivasan, Xiang Li

Speaker diarization (SD) is typically used with an automatic speech recognition (ASR) system to ascribe speaker labels to recognized words. The conventional approach reconciles outputs from independently optimized ASR an…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Speaker Tagging Correction With Non-Autoregressive Language Models

2024-08-30 · Grigor Kirakosyan, Davit Karamyan

Speech applications dealing with conversations require not only recognizing the spoken words but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+4

DiaCorrect: Error Correction Back-end For Speaker Diarization

2023-09-15 · Jiangyu Han, Federico Landini, Johan Rohdin, Mireia Diez 외

In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic sp…

Automatic Speech RecognitionDecoderspeaker-diarizationSpeaker Diarization+2

Computer-assisted Speaker Diarization: How to Evaluate Human Corrections

2018-05-01 · LREC 2018 5 · Pierre-Alex Broux, re, David Doukhan, Simon Petitrenaud 외
Active LearningFace RecognitionOptical Character Recognition (OCR)speaker-diarization+3