paper-with-me

Papers

Unsupervised Language agnostic WER Standardization

2023-03-09 · Satarupa Guha, Rahul Ambavat, Ankur Gupta, Manish Gupta, Rupeshkumar Mehta

Word error rate (WER) is a standard metric for the evaluation of Automated Speech Recognition (ASR) systems. However, WER fails to provide a fair evaluation of human perceived quality in presence of spelling variations, abbreviations, or compound words arising out of agglutination. Multiple spelling variations might be acceptable based on locale/geography, alternative abbreviations, borrowed words, and transliteration of code-mixed words from a foreign language to the target language script. Similarly, in case of agglutination, often times the agglutinated, as well as the split forms, are acceptable. Previous work handled this problem by using manually identified normalization pairs and applying them to both the transcription and the hypothesis before computing WER. In this paper, we propose an automatic WER normalization system consisting of two modules: spelling normalization and segmentation normalization. The proposed system is unsupervised and language agnostic, and therefore scalable. Experiments with ASR on 35K utterances across four languages yielded an average WER reduction of 13.28%. Human judgements of these automatically identified normalization pairs show that our WER-normalized evaluation is highly consistent with the perceived quality of ASR output.

📄 PDF Abstract BibTeX arXiv:2303.05046

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionTransliteration

Similar Papers 제목 키워드 기반

LLM4Jobs: Unsupervised occupation extraction and standardization leveraging Large Language Models

2023-09-18 · Nan Li, Bo Kang, Tijl De Bie

Automated occupation extraction and standardization from free-text job postings and resumes are crucial for applications like job recommendation and labor market policy formation. This paper introduces LLM4Jobs, a novel …

Natural Language Understanding

RetSTA: An LLM-Based Approach for Standardizing Clinical Fundus Image Reports

2025-03-12 · Jiushen Cai, Weihang Zhang, Hanruo Liu, Ningli Wang 외

Standardization of clinical reports is crucial for improving the quality of healthcare and facilitating data integration. The lack of unified standards, including format, terminology, and style, is a great challenge in c…

Data IntegrationDiagnostic

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

2026-07-06 · Xin Chen, Dongliang Xu, Cunhao Zhu, Xudong Luo 외 arxiv

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnostic ability over given medical images and texts, implicitly assuming that standardized …

Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System

2025-08-17 · Eranga Bandara, Ross Gore, Sachin Shetty, Ravi Mukkamala 외 arxiv

Accurate assessment of neuromuscular reflexes, such as the H-reflex, plays a critical role in sports science, rehabilitation, and clinical neurology. Traditional analysis of H-reflex EMG waveforms is subject to variabili…

Prompt Engineering

Harmonizing Flows: Leveraging normalizing flows for unsupervised and source-free MRI harmonization

2024-07-22 · Farzad Beizaee, Gregory A. Lodygensky, Chris L. Adamson, Deanne K. Thompso 외

Lack of standardization and various intrinsic parameters for magnetic resonance (MR) image acquisition results in heterogeneous images across different sites and devices, which adversely affects the generalization of dee…

Age EstimationMRI segmentation