paper-with-me

홈 › Papers

Investigating Diatopic Variation in a Historical Corpus

2017-04-01 · WS 2017 4 · Stefanie Dipper, S Waldenberger, ra

This paper investigates diatopic variation in a historical corpus of German. Based on equivalent word forms from different language areas, replacement rules and mappings are derived which describe the relations between these word forms. These rules and mappings are then interpreted as reflections of morphological, phonological or graphemic variation. Based on sample rules and mappings, we show that our approach can replicate results from historical linguistics. While previous studies were restricted to predefined word lists, or confined to single authors or texts, our approach uses a much wider range of data available in historical corpora.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

``Voices of the Great War'': A Richly Annotated Corpus of Italian Texts on the First World War

2020-05-01 · LREC 2020 5 · Federico Boschetti, Irene De Felice, Stefano Dei Rossi, Felice Dell{'}Orletta 외

{``}Voices of the Great War{''} is the first large corpus of Italian historical texts dating back to the period of First World War. This corpus differs from other existing resources in several respects. First, from the l…

Mapping Diatopic and Diachronic Variation in Spoken Czech: The ORTOFON and DIALEKT Corpora

2014-05-01 · LREC 2014 5 · Marie Kop{\v{r}}ivov{\'a}, Hana Gol{\'a}{\v{n}}ov{\'a}, Petra Klime{\v{s}}ov{\'a}, David Luke{\v{s}}

ORTOFON and DIALEKT are two corpora of spoken Czech (recordings + transcripts) which are currently being built at the Institute of the Czech National Corpus. The first one (ORTOFON) continues the tradition of the CNC{'}s…

Speech Recognition

DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish

2022-06-01 · LREC 2022 6 · Gioia Baldissin, Dominik Schlechtweg, Sabine Schulte im Walde

We provide a novel dataset – DiaWUG – with judgements on diatopic lexical semantic variation for six Spanish variants in Europe and Latin America. In contrast to most previous meaning-based resources and studies on seman…

OcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan

2022-10-01 · VarDial (COLING) 2022 10 · Aleksandra Miletic, Yves Scherrer

This paper presents OcWikiDisc, a new freely available corpus in Occitan, as well as language identification experiments on Occitan done as part of the corpus building process. Occitan is a regional language spoken mainl…

8kLanguage Identification

VOLIP: a corpus of spoken Italian and a virtuous example of reuse of linguistic resources

2014-05-01 · LREC 2014 5 · Iol Alfano, a, Francesco Cutugno, Aurelio De Rosa 외

The corpus VoLIP (The Voice of LIP) is an Italian speech resource which associates the audio signals to the orthographic transcriptions of the LIP Corpus. The LIP Corpus was designed to represent diaphasic, diatopic and …