paper-with-me

홈 › Papers

Topic Identification For Spontaneous Speech: Enriching Audio Features With Embedded Linguistic Information

2023-07-21 · Dejan Porjazovski, Tamás Grósz, Mikko Kurimo

Traditional topic identification solutions from audio rely on an automatic speech recognition system (ASR) to produce transcripts used as input to a text-based model. These approaches work well in high-resource scenarios, where there are sufficient data to train both components of the pipeline. However, in low-resource situations, the ASR system, even if available, produces low-quality transcripts, leading to a bad text-based classifier. Moreover, spontaneous speech containing hesitations can further degrade the performance of the ASR model. In this paper, we investigate alternatives to the standard text-only solutions by comparing audio-only and hybrid techniques of jointly utilising text and audio features. The models evaluated on spontaneous Finnish speech demonstrate that purely audio-based solutions are a viable option when ASR components are not available, while the hybrid multi-modal solutions achieve the best results.

📄 PDF Abstract BibTeX arXiv:2307.11450

Code (1)

aalto-speech/Topic-identification-for-spontaneous-Finnish-speech 공식 구현 pytorch

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay

2020-05-01 · LREC 2020 5 · Yuki Saito, Shinnosuke Takamichi, Hiroshi Saruwatari

Developing a spontaneous speech corpus would be beneficial for spoken language processing and understanding. We present a speech corpus named the SMASH corpus, which includes spontaneous speech of two Japanese male comme…

A Novel Scheme to classify Read and Spontaneous Speech

2023-06-13 · Sunil Kumar Kopparapu

The COVID-19 pandemic has led to an increased use of remote telephonic interviews, making it important to distinguish between scripted and spontaneous speech in audio recordings. In this paper, we propose a novel scheme …

SwissGPC v1.0 -- The Swiss German Podcasts Corpus

2025-09-24 · Samuel Stucki, Mark Cieliebak, Jan Deriu arxiv

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to ta…

GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

2021-06-13 · Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du 외

This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for…

Sentencespeech-recognitionSpeech Recognition

SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis

2024-08-13 · Osamu Take, Shinnosuke Takamichi, Kentaro Seki, Yoshiaki Bando 외

This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environmental factors and then control the feat…

Speech SynthesisSpoken Dialogue Systemstext-to-speechText to Speech