paper-with-me

홈 › Papers

DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization

2026-09-15 · Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante, João Miguel Sanches arxiv

Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for controlled two-party evaluation, and PDCH-HAMD, comprising voice-converted real Chinese interviews for cross-lingual validation. Cascaded systems combine speaker diarization with role-assignment heuristics, so errors can propagate across stages. We propose an end-to-end model, which we named DiaWhisper, that fine-tunes Whisper-large-v3 with LoRA and an auxiliary frame-level role head for transcription and attribution, together with DiaWhisper-DPO, a failure-mined refinement that uses genuine decoding failures as DPO rejected completions without human preference annotation. On 29 DAIC-WOZ test sessions, DiaWhisper-DPO achieves 0.973 role accuracy and 0.119 DER, 72% below the strongest cascaded baseline, and reduces seed variation from σ = .205 to .002. Retrained on PDCH-HAMD, it achieves 0.757 role accuracy and improves all 78 session-seed pairs.

📄 PDF Abstract BibTeX arXiv:2609.16661

Code (3)

InsomaniacElf/sg-tamil-tts-resources- ★ 1
Tavish9/awesome-daily-AI-arxiv ★ 117
grrlkk/writing-agent-arxiv-daily

Tasks

Speaker Diarization

Similar Papers 제목 키워드 기반

Human and Automatic Speech Recognition Performance on German Oral History Interviews

2022-01-18 · Michael Gref, Nike Matthiesen, Christoph Schmidt, Sven Behnke 외

Automatic speech recognition systems have accomplished remarkable improvements in transcription accuracy in recent years. On some domains, models now achieve near-human performance. However, transcription performance on …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Towards Automatic Transcription of ILSE ― an Interdisciplinary Longitudinal Study of Adult Development and Aging

2016-05-01 · LREC 2016 5 · Jochen Weiner, Claudia Frankenberg, Dominic Telaar, Britta Wendelstein 외

The Interdisciplinary Longitudinal Study on Adult Development and Aging (ILSE) was created to facilitate the study of challenges posed by rapidly aging societies in developed countries such as Germany. ILSE contains over…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews

2023-07-03 · Sergio Burdisso, Esaú Villatoro-Tello, Srikanth Madikeri, Petr Motlicek

We propose a simple approach for weighting self-connecting edges in a Graph Convolutional Network (GCN) and show its impact on depression detection from transcribed clinical interviews. To this end, we use a GCN for mode…

Depression Detection

Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research

2018-05-01 · LREC 2018 5 · Michael Gref, Joachim K{\"o}hler, Almut Leh
Automatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Recognition

Clinical Dialogue Transcription Error Correction using Seq2Seq Models

2022-05-26 · Gayani Nanayakkara, Nirmalie Wiratunga, David Corsar, Kyle Martin 외

Good communication is critical to good healthcare. Clinical dialogue is a conversation between health practitioners and their patients, with the explicit goal of obtaining and sharing medical information. This informatio…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decision Makingspeech-recognition+2