paper-with-me

홈 › Papers

Coconstructions in spoken data: UD annotation guidelines and first results

2026-03-30 · Ludovica Pannitto, Sylvain Kahane, Kaja Dobrovoljc, Elena Battaglia, Bruno Guillaume, Caterina Mauri, Eleonora Zucchini arxiv

The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backchannels - in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.

📄 PDF Abstract BibTeX arXiv:2603.28261

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Universal Dependencies for Western Sierra Puebla Nahuatl

2022-06-01 · LREC 2022 6 · Robert Pugh, Marivel Huerta Mendez, Mitsuya Sasaki, Francis Tyers

We present a morpho-syntactically-annotated corpus of Western Sierra Puebla Nahuatl that conforms to the annotation guidelines of the Universal Dependencies project. We describe the sources of the texts that make up the …

Annotation guidelines of UD and SUD treebanks for spoken corpora: A proposal

2021-12-01 · ACL (TLT, SyntaxFest) 2021 12 · Sylvain Kahane, Bernard Caron, Emmett Strickland, Kim Gerdes

FOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German

2016-05-01 · LREC 2016 5 · Swantje Westpfahl, Thomas Schmidt

In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard o…

LemmatizationPart-Of-Speech TaggingPOS

Spoken Language Treebanks in Universal Dependencies: an Overview

2022-06-01 · LREC 2022 6 · Kaja Dobrovoljc

Given the benefits of syntactically annotated collections of transcribed speech in spoken language research and applications, many spoken language treebanks have been developed in the last decades, with divergent annotat…

ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

2024-03-27 · Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)LemmatizationPart-Of-Speech Tagging+2