paper-with-me

홈 › Papers

Iterative Language Model Adaptation for Indo-Aryan Language Identification

2018-08-01 · COLING 2018 8 · Tommi Jauhiainen, Heidi Jauhiainen, Krister Lind{\'e}n

This paper presents the experiments and results obtained by the SUKI team in the Indo-Aryan Language Identification shared task of the VarDial 2018 Evaluation Campaign. The shared task was an open one, but we did not use any corpora other than what was distributed by the organizers. A total of eight teams provided results for this shared task. Our submission using a HeLI-method based language identifier with iterative language model adaptation obtained the best results in the shared task with a macro F1-score of 0.958.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Toward a deep dialectological representation of Indo-Aryan

2019-06-01 · WS 2019 6 · Chundra Cathcart

This paper presents a new approach to disentangling inter-dialectal and intra-dialectal relationships within one such group, the Indo-Aryan subgroup of Indo-European. We draw upon admixture models and deep generative mod…

IndoNLP 2025: Shared Task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages

2025-01-10 · Deshan Sumanathilaka, Isuri Anuradha, Ruvan Weerasinghe, Nicholas Micallef 외

The paper overviews the shared task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages. It focuses on the reverse transliteration of low-resourced languages in the Indo-Aryan family to their native s…

Transliteration

Complexity counts: global and local perspectives on Indo-Aryan numeral systems

2025-05-19 · Chundra Cathcart

The numeral systems of Indo-Aryan languages such as Hindi, Gujarati, and Bengali are highly unusual in that unlike most numeral systems (e.g., those of English, Chinese, etc.), forms referring to 1--99 are highly non-tra…

A probabilistic assessment of the Indo-Aryan Inner-Outer Hypothesis

2019-11-29 · Chundra A. Cathcart

This paper uses a novel data-driven probabilistic approach to address the century-old Inner-Outer hypothesis of Indo-Aryan. I develop a Bayesian hierarchical mixed-membership model to assess the validity of this hypothes…

Discriminating between Indo-Aryan Languages Using SVM Ensembles

2018-07-09 · COLING 2018 8 · Alina Maria Ciobanu, Marcos Zampieri, Shervin Malmasi, Santanu Pal 외

In this paper we present a system based on SVM ensembles trained on characters and words to discriminate between five similar languages of the Indo-Aryan family: Hindi, Braj Bhasha, Awadhi, Bhojpuri, and Magahi. We inves…

Language Identification