Language Identification
6개 벤치마크 · 논문 851편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
Scaling Speech Technology to 1,000+ Languages
SpeechBrain: A General-Purpose Speech Toolkit
The WiLI benchmark dataset for written language identification
GlotLID: Language Identification for Low-Resource Languages
Papers
Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline t…
Language IdentificationImproving Language Identification for Code-Switched Utterances with Integer Linear Programming
Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this pa…
Language IdentificationReference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO)…
Language IdentificationReinforcement LearningMachine TranslationLower-Resource, Higher Scores: Language Bias in LLM Evaluators
LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, l…
Language IdentificationLanguage Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
Language identification is commonly addressed using either neural architectures or statistical n-gram models. Neural approaches typically require substantial computational resources, whereas classical frequency-based met…
Language IdentificationLanguage Identification with Succinct Machine-Independent Traces
Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negativ…
Language Identification