paper-with-me

Papers

Finding Variants for Construction-Based Dialectometry: A Corpus-Based Approach to Regional CxGs

2021-04-03 · Jonathan Dunn

This paper develops a construction-based dialectometry capable of identifying previously unknown constructions and measuring the degree to which a given construction is subject to regional variation. The central idea is to learn a grammar of constructions (a CxG) using construction grammar induction and then to use these constructions as features for dialectometry. This offers a method for measuring the aggregate similarity between regional CxGs without limiting in advance the set of constructions subject to variation. The learned CxG is evaluated on how well it describes held-out test corpora while dialectometry is evaluated on how well it can model regional varieties of English. Themethod is tested using two distinct datasets: First, the International Corpus of English representing eight outer circle varieties; Second, a web-crawled corpus representing five inner circle varieties. Results show that themethod (1) produces a grammar with stable quality across sub-sets of a single corpus that is (2) capable of distinguishing between regional varieties of Englishwith a high degree of accuracy, thus (3) supporting dialectometricmethods formeasuring the similarity between varieties of English and (4) measuring the degree to which each construction is subject to regional variation. This is important for cognitive sociolinguistics because it operationalizes the idea that competition between constructions is organized at the functional level so that dialectometry needs to represent as much of the available functional space as possible.

📄 PDF Abstract BibTeX arXiv:2104.01299

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Global Syntactic Variation in Seven Languages: Towards a Computational Dialectology

2021-04-03 · Jonathan Dunn

The goal of this paper is to provide a complete representation of regional linguistic variation on a global scale. To this end, the paper focuses on removing three constraints that have previously limited work within dia…

Tools for Building a Corpus to Study the Historical and Geographical Variation of the Romanian Language

2017-09-01 · RANLP 2017 9 · Victoria Bobicev, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Cenel Augusto Perez

Contemporary standard language corpora are ideal for NLP. There are few morphologically and syntactically annotated corpora for Romanian, and those existing or in progress only deal with the Contemporary Romanian standar…

Multi-dialect Neural Machine Translation and Dialectometry

2018-12-01 · PACLIC 2018 12 · Kaori Abe, Yuichiroh Matsubayashi, Naoaki Okazaki, Kentaro Inui
Machine TranslationTranslation

dialectR: Doing Dialectometry in R

2022-10-01 · VarDial (COLING) 2022 10 · Ryan Soh-Eun Shim, John Nerbonne

We present dialectR, an open-source R package for performing quantitative analyses of dialects based on categorical measures of difference and on variants of edit distance. dialectR stands as one of the first programmabl…

Clustering

Learning about Spanish dialects through Twitter

2015-11-16 · Bruno Gonçalves, David Sánchez

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. Th…

BIG-bench Machine Learning