paper-with-me

Papers

Comparing Human and Machine Errors in Conversational Speech Transcription

2017-08-29 · Andreas Stolcke, Jasha Droppo

Recent work in automatic recognition of conversational telephone speech (CTS) has achieved accuracy levels comparable to human transcribers, although there is some debate how to precisely quantify human performance on this task, using the NIST 2000 CTS evaluation set. This raises the question what systematic differences, if any, may be found differentiating human from machine transcription errors. In this paper we approach this question by comparing the output of our most accurate CTS recognition system to that of a standard speech transcription vendor pipeline. We find that the most frequent substitution, deletion and insertion error types of both outputs show a high degree of overlap. The only notable exception is that the automatic recognizer tends to confuse filled pauses ("uh") and backchannel acknowledgments ("uhhuh"). Humans tend not to make this error, presumably due to the distinctive and opposing pragmatic functions attached to these words. Furthermore, we quantify the correlation between human and machine errors at the speaker level, and investigate the effect of speaker overlap between training and test data. Finally, we report on an informal "Turing test" asking humans to discriminate between automatic and human transcription error cases.

📄 PDF Abstract BibTeX arXiv:1708.08615

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Disfluencies and Human Speech Transcription Errors

2019-04-08 · Vicky Zayats, Trang Tran, Richard Wright, Courtney Mansfield 외

This paper explores contexts associated with errors in transcrip-tion of spontaneous speech, shedding light on human perceptionof disfluencies and other conversational speech phenomena. Anew version of the Switchboard co…

ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations

2025-07-17 · Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira 외 arxiv

The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user inten…

Improving endpoint detection in end-to-end streaming ASR for conversational speech

2025-05-19 · Anandh C, Karthik Pandia Durai, Jeena Prakash, Manickavela Arumugam 외

ASR endpointing (EP) plays a major role in delivering a good user experience in products supporting human or artificial agents in human-human/machine conversations. Transducer-based ASR (T-ASR) is an end-to-end (E2E) ASR…

Action DetectionActivity Detection

Cross-Lingual Conversational Speech Summarization with Large Language Models

2024-08-12 · Max Nelson, Shannon Wotherspoon, Francis Keith, William Hartmann 외

Cross-lingual conversational speech summarization is an important problem, but suffers from a dearth of resources. While transcriptions exist for a number of languages, translated conversational speech is rare and datase…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Evaluation of Automated Speech Recognition Systems for Conversational Speech: A Linguistic Perspective

2022-11-05 · Hannaneh B. Pasandi, Haniyeh B. Pasandi

Automatic speech recognition (ASR) meets more informal and free-form input data as voice user interfaces and conversational agents such as the voice assistants such as Alexa, Google Home, etc., gain popularity. Conversat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition