paper-with-me

홈 › Papers

Building a Time-Aligned Cross-Linguistic Reference Corpus from Language Documentation Data (DoReCo)

2020-05-01 · LREC 2020 5 · Ludger Paschen, Fran{\c{c}}ois Delafontaine, Christoph Draxler, Susanne Fuchs, Matthew Stave, Frank Seifart

Natural speech data on many languages have been collected by language documentation projects aiming to preserve lingustic and cultural traditions in audivisual records. These data hold great potential for large-scale cross-linguistic research into phonetics and language processing. Major obstacles to utilizing such data for typological studies include the non-homogenous nature of file formats and annotation conventions found both across and within archived collections. Moreover, time-aligned audio transcriptions are typically only available at the level of broad (multi-word) phrases but not at the word and segment levels. We report on solutions developed for these issues within the DoReCo (DOcumentation REference COrpus) project. DoReCo aims at providing time-aligned transcriptions for at least 50 collections of under-resourced languages. This paper gives a preliminary overview of the current state of the project and details our workflow, in particular standardization of formats and conventions, the addition of segmental alignments with WebMAUS, and DoReCo{'}s applicability for subsequent research programs. By making the data accessible to the scientific community, DoReCo is designed to bridge the gap between language documentation and linguistic inquiry.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PAWS: A Multi-lingual Parallel Treebank with Anaphoric Relations

2018-06-01 · WS 2018 6 · Anna Nedoluzhko, Michal Nov{\'a}k, Maciej Ogrodniczuk

We present PAWS, a multi-lingual parallel treebank with coreference annotation. It consists of English texts from the Wall Street Journal translated into Czech, Russian and Polish. In addition, the texts are syntacticall…

Coreference ResolutionMachine Translation

Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization

2024-12-19 · Sahil Wadhwa, Chengtian Xu, Haoming Chen, Aakash Mahalingam 외

The automatic generation of counter-speech (CS) is a critical strategy for addressing hate speech by providing constructive and informed responses. However, existing methods often fail to generate high-quality, impactful…

Building a Manually Annotated Hungarian Coreference Corpus: Workflow and Tools

2022-10-01 · COLING (CRAC) 2022 10 · Noémi Vadász

This paper presents the complete workflow of building a manually annotated Hungarian corpus, KorKor, with particular reference to anaphora and coreference annotation. All linguistic annotation layers were corrected manua…

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

2026-04-14 · Qixi Zheng, Yuxiang Zhao, Tianrui Wang, Wenxi Chen 외 arxiv

Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building ze…

Voice Conversion

Using Linguistic Features to Improve the Generalization Capability of Neural Coreference Resolvers

2017-08-01 · EMNLP 2018 10 · Nafise Sadat Moosavi, Michael Strube

Coreference resolution is an intermediate step for text understanding. It is used in tasks and domains for which we do not necessarily have coreference annotated corpora. Therefore, generalization is of special importanc…

coreference-resolutionCoreference Resolution