paper-with-me

Papers

Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication

2024-07-11 · Arif Ahmad, Mothika Gayathri Khyathi, Pushpak Bhattacharyya

Reduplication and repetition, though similar in form, serve distinct linguistic purposes. Reduplication is a deliberate morphological process used to express grammatical, semantic, or pragmatic nuances, while repetition is often unintentional and indicative of disfluency. This paper presents the first large-scale study of reduplication and repetition in speech using computational linguistics. We introduce IndicRedRep, a new publicly available dataset containing Hindi, Telugu, and Marathi text annotated with reduplication and repetition at the word level. We evaluate transformer-based models for multi-class reduplication and repetition token classification, utilizing the Reparandum-Interregnum-Repair structure to distinguish between the two phenomena. Our models achieve macro F1 scores of up to 85.62% in Hindi, 83.95% in Telugu, and 84.82% in Marathi for reduplication-repetition classification.

📄 PDF Abstract BibTeX arXiv:2407.08147

Code (0)

등록된 구현이 없습니다.

Tasks

token-classificationToken Classification

Similar Papers 제목 키워드 기반

Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis

2025-11-17 · Zaara Zabeen Arpa, Sadnam Sakib Apurbo, Nazia Karim Khan Oishee, Ajwad Abrar arxiv

Automatic Speech Recognition (ASR) transcripts, especially in low-resource languages like Bangla, contain a critical ambiguity: word-word repetitions can be either Repetition Disfluency (unintentional ASR error/hesitatio…

Speech Recognition

A self-reliant finite automata for reduplication detection

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Reduplication is a common phenomenon in almost all human languages. It implies the repetition of the smallest linguistic unit partially (e.g. flip flop) or completely (e.g. bye bye). Symbolically it can be written as WiW…

Typical vs. Atypical Disfluency Classification: Introducing the IIITH-TISA Corpus and Temporal Context-Based Feature Representations

2024-11-26 · Priyanka Kommagouni, Vamshiraghusimha Narasinga, Purva Barche, Sai Akarsh C 외

Speech disfluencies in spontaneous communication can be categorized as either typical or atypical. Typical disfluencies, such as hesitations and repetitions, are natural occurrences in everyday speech, while atypical dis…

Detecting Reduplication in Videos of American Sign Language

2012-05-01 · LREC 2012 5 · Zoya Gavrilov, Stan Sclaroff, Carol Neidle, Sven Dickinson

A framework is proposed for the detection of reduplication in digital videos of American Sign Language (ASL). In ASL, reduplication is used for a variety of linguistic purposes, including overt marking of plurality on no…

Planning and Generating Natural and Diverse Disfluent Texts as Augmentation for Disfluency Detection

2020-11-01 · EMNLP 2020 11 · Jingfeng Yang, Diyi Yang, Zhaoran Ma

Existing approaches to disfluency detection heavily depend on human-annotated data. Numbers of data augmentation methods have been proposed to alleviate the dependence on labeled data. However, current augmentation appro…

Data Augmentation