paper-with-me

홈 › Papers

HeLI-OTS, Off-the-shelf Language Identifier for Text

2022-06-01 · LREC 2022 6 · Tommi Jauhiainen, Heidi Jauhiainen, Krister Lindén

This paper introduces HeLI-OTS, an off-the-shelf text language identification tool using the HeLI language identification method. The HeLI-OTS language identifier is equipped with language models for 200 languages and licensed for academic as well as commercial use. We present the HeLI method and its use in our previous research. Then we compare the performance of the HeLI-OTS language identifier with that of fastText on two different data sets, showing that fastText favors the recall of common languages, whereas HeLI-OTS reaches both high recall and high precision for all languages. While introducing existing off-the-shelf language identification tools, we also give a picture of digital humanities-related research that uses such tools. The validity of the results of such research depends on the results given by the language identifier used, and especially for research focusing on the less common languages, the tendency to favor widely used languages might be very detrimental, which Heli-OTS is now able to remedy.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

HeLI-based Experiments in Discriminating Between Dutch and Flemish Subtitles

2018-08-01 · COLING 2018 8 · Tommi Jauhiainen, Heidi Jauhiainen, Krister Lind{\'e}n

This paper presents the experiments and results obtained by the SUKI team in the Discriminating between Dutch and Flemish in Subtitles shared task of the VarDial 2018 Evaluation Campaign. Our best submission was ranked 8…

ClusteringLanguage IdentificationText Categorization

Iterative Language Model Adaptation for Indo-Aryan Language Identification

2018-08-01 · COLING 2018 8 · Tommi Jauhiainen, Heidi Jauhiainen, Krister Lind{\'e}n

This paper presents the experiments and results obtained by the SUKI team in the Indo-Aryan Language Identification shared task of the VarDial 2018 Evaluation Campaign. The shared task was an open one, but we did not use…

Language IdentificationLanguage ModelingLanguage Modelling

SCALAR: A Part-of-speech Tagger for Identifiers

2025-04-23 · Christian D. Newman, Brandon Scholten, Sophia Testa, Joshua A. C. Behler 외

The paper presents the Source Code Analysis and Lexical Annotation Runtime (SCALAR), a tool specialized for mapping (annotating) source code identifier names to their corresponding part-of-speech tag sequence (grammar pa…

TAG

Language Identification and Named Entity Recognition in Hinglish Code Mixed Tweets

2018-07-01 · ACL 2018 7 · Kushagra Singh, Indira Sen, Ponnurangam Kumaraguru

While growing code-mixed content on Online Social Networks(OSN) provides a fertile ground for studying various aspects of code-mixing, the lack of automated text analysis tools render such studies challenging. To meet th…

Abuse DetectionChunkingLanguage Identificationnamed-entity-recognition+5

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

2026-08-01 · Bogdan Savelyev arxiv

Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set wh…