paper-with-me

Papers

A novel multimodal dynamic fusion network for disfluency detection in spoken utterances

2022-11-27 · Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Manan Suri, Rajiv Ratn Shah

Disfluency, though originating from human spoken utterances, is primarily studied as a uni-modal text-based Natural Language Processing (NLP) task. Based on early-fusion and self-attention-based multimodal interaction between text and acoustic modalities, in this paper, we propose a novel multimodal architecture for disfluency detection from individual utterances. Our architecture leverages a multimodal dynamic fusion network that adds minimal parameters over an existing text encoder commonly used in prior art to leverage the prosodic and acoustic cues hidden in speech. Through experiments, we show that our proposed model achieves state-of-the-art results on the widely used English Switchboard for disfluency detection and outperforms prior unimodal and multimodal systems in literature by a significant margin. In addition, we make a thorough qualitative analysis and show that, unlike text-only systems, which suffer from spurious correlations in the data, our system overcomes this problem through additional cues from speech signals. We make all our codes publicly available on GitHub.

📄 PDF Abstract BibTeX arXiv:2211.14700

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal interaction

Similar Papers 제목 키워드 기반

Incremental Disfluency Detection for Spoken Learner English

2022-07-01 · NAACL (BEA) 2022 7 · Lucy Skidmore, Roger Moore

Incremental disfluency detection provides a framework for computing communicative meaning from hesitations, repetitions and false starts commonly found in speech. One application of this area of research is in dialogue-b…

Missingness-resilient Video-enhanced Multimodal Disfluency Detection

2024-06-11 · Payal Mohapatra, Shamika Likhite, Subrata Biswas, Bashima Islam 외

Most existing speech disfluency detection techniques only rely upon acoustic data. In this work, we present a practical multimodal disfluency detection approach that leverages available video data together with audio. We…

What BERT Based Language Models Learn in Spoken Transcripts: An Empirical Study

2021-09-19 · Ayush Kumar, Mukuntha Narayanan Sundararaman, Jithendra Vepa

Language Models (LMs) have been ubiquitously leveraged in various tasks including spoken language understanding (SLU). Spoken language requires careful understanding of speaker interactions, dialog states and speech indu…

Spoken Language Understanding

What BERT Based Language Model Learns in Spoken Transcripts: An Empirical Study

2021-11-01 · EMNLP (BlackboxNLP) 2021 11 · Ayush Kumar, Mukuntha Narayanan Sundararaman, Jithendra Vepa

Language Models (LMs) have been ubiquitously leveraged in various tasks including spoken language understanding (SLU). Spoken language requires careful understanding of speaker interactions, dialog states and speech indu…

Language ModelingLanguage ModellingSpoken Language Understanding

Span Classification with Structured Information for Disfluency Detection in Spoken Utterances

2022-03-30 · Sreyan Ghosh, Sonal Kumar, Yaman Kumar Singla, Rajiv Ratn Shah 외

Existing approaches in disfluency detection focus on solving a token-level classification task for identifying and removing disfluencies in text. Moreover, most works focus on leveraging only contextual information captu…

Classification