paper-with-me

Papers

CLD²: Language Documentation Meets Natural Language Processing for Revitalising Endangered Languages

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Language revitalisation should not be understood as a direct outcome of language documentation, which is mainly focused on the creation of language repositories. Natural language processing (NLP) offers the potential to complement and exploit these repositories through the development of language technologies that may directly impact the vitality status of endangered languages. In this paper, we discuss the current state of the interaction between language documentation and computational linguistics, present a diagnosis of how the outputs of recent documentation projects for endangered languages are underutilised for the NLP community, and discuss how the situation could change from both the documentary linguistics and NLP perspectives. All this is introduced as a bridging paradigm called Computational Language Documentation and Development (CLD²). CLD² calls for (1) the inclusion of NLP-friendly annotated data as a deliverable of future language documentation projects; and (2) the exploitation of language documentation databases by the NLP community to promote the computerisation of endangered languages at a global scale.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CLD² Language Documentation Meets Natural Language Processing for Revitalising Endangered Languages

2022-05-01 · ComputEL (ACL) 2022 5 · Roberto Zariquiey, Arturo Oncevay, Javier Vera

Language revitalisation should not be understood as a direct outcome of language documentation, which is mainly focused on the creation of language repositories. Natural language processing (NLP) offers the potential to …

Proceedings of the 2017 EMNLP Workshop: Natural Language Processing meets Journalism

2017-09-01 · WS 2017 9 ·

Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards

2021-08-16 · ACL (GEM) 2021 8 · Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez, Pawan Sasanka Ammanamanchi 외

Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of the people involved in the building of n…

Text Generation

A parallel corpus of Python functions and documentation strings for automated code documentation and code generation

2017-07-07 · IJCNLP 2017 11 · Antonio Valerio Miceli Barone, Rico Sennrich

Automated documentation of programming source code and automated code generation from natural language are challenging tasks of both practical and scientific interest. Progress in these areas has been limited by the low …

Code GenerationData AugmentationMachine TranslationTranslation

Towards a General-Purpose Linguistic Annotation Backend

2018-12-13 · Graham Neubig, Patrick Littell, Chian-Yu Chen, Jean Lee 외

Language documentation is inherently a time-intensive process; transcription, glossing, and corpus management consume a significant portion of documentary linguists' work. Advances in natural language processing can help…

Management