paper-with-me

Papers

Crowdsourced and Automatic Speech Prominence Estimation

2023-10-12 · Max Morrison, Pranav Pawar, Nathan Pruyne, Jennifer Cole, Bryan Pardo

The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the prominence of each word in an utterance. These prominence labels are useful for linguistic analysis, as well as training automated systems to perform emphasis-controlled text-to-speech or emotion recognition. Manually annotating prominence is time-consuming and expensive, which motivates the development of automated methods for speech prominence estimation. However, developing such an automated system using machine-learning methods requires human-annotated training data. Using our system for acquiring such human annotations, we collect and open-source crowdsourced annotations of a portion of the LibriTTS dataset. We use these annotations as ground truth to train a neural speech prominence estimator that generalizes to unseen speakers, datasets, and speaking styles. We investigate design decisions for neural prominence estimation as well as how neural prominence estimation improves as a function of two key factors of annotation cost: dataset size and the number of annotations per utterance.

📄 PDF Abstract BibTeX arXiv:2310.08464

Code (1)

reseval/reseval 공식 구현

Tasks

Emotion Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Prominence-aware automatic speech recognition for conversational speech

2025-09-12 · Julian Linke, Barbara Schuppler arxiv

This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-…

Speech Recognition

A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings

2024-12-11 · Anindita Mondal, Rangavajjala Sankara Bharadwaj, Jhansi Mallela, Anil Kumar Vuppala 외

Automatic detection of prominence at the word and syllable-levels is critical for building computer-assisted language learning systems. It has been shown that prosody embeddings learned by the current state-of-the-art (S…

text-to-speechText to Speech

Human Transcription Quality Improvement

2023-09-24 · Jian Gao, Hanbo Sun, Cheng Cao, Zheng Du

High quality transcription data is crucial for training automatic speech recognition (ASR) systems. However, the existing industry-level data collection pipelines are expensive to researchers, while the quality of crowds…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

2026-05-14 · Ivan Molodetskikh, Kirill Malyshev, Mark Mirgaleev, Nikita Zagainov 외 arxiv

Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and texture distortions that degrade perceived quality. These defects var…

Image Super-Resolution

Prominence-Aware Artifact Detection and Dataset for Image Super-Resolution

2025-10-19 · Ivan Molodetskikh, Kirill Malyshev, Mark Mirgaleev, Nikita Zagainov 외 arxiv

Generative single-image super-resolution (SISR) is advancing rapidly, yet even state-of-the-art models produce visual artifacts: unnatural patterns and texture distortions that degrade perceived quality. These defects va…

Image Super-ResolutionArtifact Detection