paper-with-me

Papers

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

2026-04-30 · Thibault Bañeras Roux, Jane Wottawa, Mickael Rouvier, Teva Merlin, Richard Dufour arxiv

Conventionally, Automatic Speech Recognition (ASR) systems are evaluated on their ability to correctly recognize each word contained in a speech signal. In this context, the word error rate (WER) metric is the reference for evaluating speech transcripts. Several studies have shown that this measure is too limited to correctly evaluate an ASR system, which has led to the proposal of other variants of metrics (weighted WER, BERTscore, semantic distance, etc.). However, they remain system-oriented, even when transcripts are intended for humans. In this paper, we firstly present Human Assessed Transcription Side-by-side (HATS), an original French manually annotated data set in terms of human perception of transcription errors produced by various ASR systems. 143 humans were asked to choose the best automatic transcription out of two hypotheses. We investigated the relationship between human preferences and various ASR evaluation metrics, including lexical and embedding-based ones, the latter being those that correlate supposedly the most with human perception.

📄 PDF Abstract BibTeX arXiv:2604.27542

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching with Collaborative LLM-Agents

2025-03-19 · Hao Liang, Zhipeng Dong, Kaixin Chen, Jiyuan Guo 외

Surround-view perception has garnered significant attention for its ability to enhance the perception capabilities of autonomous driving vehicles through the exchange of information with surrounding cameras. However, exi…

Autonomous DrivingImage StitchingSSIM

Characterizing the public perception of WhatsApp through the lens of media

2018-08-17 · Josemar Alves Caetano, Gabriel Magno, Evandro Cunha, Wagner Meira Jr. 외

WhatsApp is, as of 2018, a significant component of the global information and communication infrastructure, especially in developing countries. However, probably due to its strong end-to-end encryption, WhatsApp became …

ArticlesMisinformation

ChatSearch: a Dataset and a Generative Retrieval Model for General Conversational Image Retrieval

2024-10-24 · Zijia Zhao, Longteng Guo, Tongtian Yue, Erdong Hu 외

In this paper, we investigate the task of general conversational image retrieval on open-domain images. The objective is to search for images based on interactive conversations between humans and computers. To advance th…

Image RetrievalRetrievalWorld Knowledge

Neural Generation Meets Real People: Building a Social, Informative Open-Domain Dialogue Agent

2022-07-25 · SIGDIAL (ACL) 2022 9 · Ethan A. Chi, Ashwin Paranjape, Abigail See, Caleb Chiam 외

We present Chirpy Cardinal, an open-domain social chatbot. Aiming to be both informative and conversational, our bot chats with users in an authentic, emotionally intelligent way. By integrating controlled neural generat…

Chatbot

Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism

2025-10-15 · Xiaoshu Chen, Sihang Zhou, Ke Liang, Duanyang Yuan 외 arxiv

Chain of thought (CoT) fine-tuning aims to endow large language models (LLMs) with reasoning capabilities by training them on curated reasoning traces. It leverages both supervised and reinforced fine-tuning to cultivate…

Mathematical ReasoningCode Generation