paper-with-me

Papers

MultiTurnCleanup: A Benchmark for Multi-Turn Spoken Conversational Transcript Cleanup

2023-05-19 · Hua Shen, Vicky Zayats, Johann C. Rocholl, Daniel D. Walker, Dirk Padfield

Current disfluency detection models focus on individual utterances each from a single speaker. However, numerous discontinuity phenomena in spoken conversational transcripts occur across multiple turns, hampering human readability and the performance of downstream NLP tasks. This study addresses these phenomena by proposing an innovative Multi-Turn Cleanup task for spoken conversational transcripts and collecting a new dataset, MultiTurnCleanup1. We design a data labeling schema to collect the high-quality dataset and provide extensive data analysis. Furthermore, we leverage two modeling approaches for experimental evaluation as benchmarks for future research.

📄 PDF Abstract BibTeX arXiv:2305.12029

Code (1)

huashen218/multiturncleanup 공식 구현

Similar Papers 제목 키워드 기반

What BERT Based Language Models Learn in Spoken Transcripts: An Empirical Study

2021-09-19 · Ayush Kumar, Mukuntha Narayanan Sundararaman, Jithendra Vepa

Language Models (LMs) have been ubiquitously leveraged in various tasks including spoken language understanding (SLU). Spoken language requires careful understanding of speaker interactions, dialog states and speech indu…

Spoken Language Understanding

What BERT Based Language Model Learns in Spoken Transcripts: An Empirical Study

2021-11-01 · EMNLP (BlackboxNLP) 2021 11 · Ayush Kumar, Mukuntha Narayanan Sundararaman, Jithendra Vepa

Language Models (LMs) have been ubiquitously leveraged in various tasks including spoken language understanding (SLU). Spoken language requires careful understanding of speaker interactions, dialog states and speech indu…

Language ModelingLanguage ModellingSpoken Language Understanding

Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling

2024-12-20 · Maximillian Chen, Ruoxi Sun, Sercan Ö. Arik

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-spec…

Multi-Task Learning

Building a Taiwanese Mandarin Spoken Language Model: A First Attempt

2024-11-11 · Chih-Kai Yang, Yu-Kuan Fu, Chen-An Li, Yi-Cheng Lin 외

This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations.…

DecoderLanguage ModelingLanguage ModellingLarge Language Model

Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

2025-03-03 · Siddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang 외

The recent wave of audio foundation models (FMs) could provide new capabilities for conversational modeling. However, there have been limited efforts to evaluate these audio FMs comprehensively on their ability to have n…

BenchmarkingSpoken Dialogue Systems