paper-with-me

Papers

Analyzing discourse functions with acoustic features and phone embeddings: non-lexical items in Taiwan Mandarin

2022-11-01 · ROCLING 2022 11 · Pin-Er Chen, Yu-Hsiang Tseng, Chi-Wei Wang, Fang-Chi Yeh, Shu-Kai Hsieh

Non-lexical items are expressive devices used in conversations that are not words but are nevertheless meaningful. These items play crucial roles, such as signaling turn-taking or marking stances in interactions. However, as the non-lexical items do not stably correspond to written or phonological forms, past studies tend to focus on studying their acoustic properties, such as pitches and durations. In this paper, we investigate the discourse functions of non-lexical items through their acoustic properties and the phone embeddings extracted from a deep learning model. Firstly, we create a non-lexical item dataset based on the interpellation video clips from Taiwan’s Legislative Yuan. Then, we manually identify the non-lexical items and their discourse functions in the videos. Next, we analyze the acoustic properties of those items through statistical modeling and building classifiers based on phone embeddings extracted from a phone recognition model. We show that (1) the discourse functions have significant effects on the acoustic features; and (2) the classifiers built on phone embeddings perform better than the ones on conventional acoustic properties. These results suggest that phone embeddings may reflect the phonetic variations crucial in differentiating the discourse functions of non-lexical items.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Discourse on ASR Measurement: Introducing the ARPOCA Assessment Tool

2022-05-01 · ACL 2022 5 · Megan Merz, Olga Scrivner

Automatic speech recognition (ASR) has evolved from a pipeline architecture with pronunciation dictionaries, phonetic features and language models to the end-to-end systems performing a direct translation from a raw wave…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

CUPE: Contextless Universal Phoneme Encoder for Language-Agnostic Speech Processing

2025-08-21 · Abdul Rehman, Jian-Jun Zhang, Xiaosong Yang arxiv

Universal phoneme recognition typically requires analyzing long speech segments and language-specific patterns. Many speech processing tasks require pure phoneme representations free from contextual influence, which moti…

PRODIS - a speech database and a phoneme-based language model for the study of predictability effects in Polish

2024-04-15 · Zofia Malisz, Jan Foremski, Małgorzata Kul

We present a speech database and a phoneme-level language model of Polish. The database and model are designed for the analysis of prosodic and discourse factors and their impact on acoustic parameters in interaction wit…

Language ModelingLanguage Modelling

The trajectoRIR Database: Room Acoustic Recordings Along a Trajectory of Moving Microphones

2025-03-29 · Stefano Damiano, Kathleen MacWilliam, Valerio Lorenzoni, Thomas Dietzen 외

Data availability is essential to develop acoustic signal processing algorithms, especially when it comes to data-driven approaches that demand large and diverse training datasets. For this reason, an increasing number o…

Sound Source Localization

Analyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech Recognition

2019-07-09 · Yonatan Belinkov, Ahmed Ali, James Glass

End-to-end neural network systems for automatic speech recognition (ASR) are trained from acoustic features to text transcriptions. In contrast to modular ASR systems, which contain separately-trained components for acou…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2