paper-with-me

홈 › Papers

NaturalTurn: A Method to Segment Transcripts into Naturalistic Conversational Turns

2024-03-22 · Gus Cooney, Andrew Reece

Conversation is the subject of increasing interest in the social, cognitive, and computational sciences. And yet, as conversational datasets continue to increase in size and complexity, researchers lack scalable methods to segment speech-to-text transcripts into conversational turns--the basic building blocks of social interaction. We introduce "NaturalTurn," a turn segmentation algorithm designed to accurately capture the dynamics of naturalistic exchange. NaturalTurn operates by distinguishing speakers' primary conversational turns from listeners' secondary utterances, such as backchannels, brief interjections, and other forms of parallel speech that characterize conversation. Using data from a large conversation corpus, we show how NaturalTurn-derived transcripts demonstrate favorable statistical and inferential characteristics compared to transcripts derived from existing methods. The NaturalTurn algorithm represents an improvement in machine-generated transcript processing methods, or "turn models" that will enable researchers to associate turn-taking dynamics with the broader outcomes that result from social interaction, a central goal of conversation science.

📄 PDF Abstract BibTeX arXiv:2403.15615

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Text

Similar Papers 제목 키워드 기반

Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models

2026-05-23 · Feng Chen, Justin Tauscher, Changye Li, Meliha Yetisgen 외 arxiv

Speech monologues recorded in naturalistic settings provide opportunities to characterize mental illness phenomenology and detect symptom exacerbation. Large language models (LLMs) offer new possibilities for automating …

MONAH: Multi-Modal Narratives for Humans to analyze conversations

2021-01-18 · EACL 2021 2 · Joshua Y. Kim, Greyson Y. Kim, Chunfeng Liu, Rafael A. Calvo 외

In conversational analyses, humans manually weave multimodal information into the transcripts, which is significantly time-consuming. We introduce a system that automatically expands the verbatim transcripts of video-rec…

Emotion Recognition in ConversationFeature EngineeringText Generation

Towards Emotion and Affect Detection in the Multimodal LAST MINUTE Corpus

2012-05-01 · LREC 2012 5 · J{\"o}rg Frommer, Bernd Michaelis, Dietmar R{\"o}sner, Andreas Wendemuth 외

The LAST MINUTE corpus comprises multimodal recordings (e.g. video, audio, transcripts) from WOZ interactions in a mundane planning task (R{\"o}sner et al., 2011). It is one of the largest corpora with naturalistic data …

MultiTurnCleanup: A Benchmark for Multi-Turn Spoken Conversational Transcript Cleanup

2023-05-19 · Hua Shen, Vicky Zayats, Johann C. Rocholl, Daniel D. Walker 외

Current disfluency detection models focus on individual utterances each from a single speaker. However, numerous discontinuity phenomena in spoken conversational transcripts occur across multiple turns, hampering human r…

What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition

2021-07-05 · Piotr Żelasko, Raghavendra Pappagari, Najim Dehak

Dialog acts can be interpreted as the atomic units of a conversation, more fine-grained than utterances, characterized by a specific communicative function. The ability to structure a conversational transcript as a seque…

SegmentationSpecificitySpoken Language Understanding