Papers Language Identification
“Language Identification” 태그가 달린 논문 851편 · 필터 해제
Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline t…
Language IdentificationImproving Language Identification for Code-Switched Utterances with Integer Linear Programming
Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this pa…
Language IdentificationReference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO)…
Language IdentificationReinforcement LearningMachine TranslationLower-Resource, Higher Scores: Language Bias in LLM Evaluators
LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, l…
Language IdentificationLanguage Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
Language identification is commonly addressed using either neural architectures or statistical n-gram models. Neural approaches typically require substantial computational resources, whereas classical frequency-based met…
Language IdentificationLanguage Identification with Succinct Machine-Independent Traces
Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negativ…
Language IdentificationFrom Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages
Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architec…
Language IdentificationSpeech RecognitionLuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish
State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, …
Language IdentificationCross-Lingual TransferActivity DetectionImproving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation
This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a…
Language IdentificationSpeech RecognitionGenerating in the Limit with Infinitely Many Hallucinations
The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language. The r…
Language IdentificationCHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios
We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language identification: cousin languages and orthographic noise. Our dataset has…
Language IdentificationThe Inclusion Depth of Pattern Languages: An Open Problem in Algorithmic Learning Theory
Pattern languages are a classical model in formal language theory and algorithmic learning theory. This note formulates the problem of computing the inclusion depth of a pattern language: the length of the longest strict…
Language IdentificationPashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A system may produce no audio, speak a neighbouring language, preserve…
Language IdentificationAudience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus
This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corpus was collected from 51,660 pages acros…
Language IdentificationMultilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unreliable in multilingual settings: most SAEs …
Language IdentificationMachine TranslationDemographic and Linguistic Bias Evaluation in Omnimodal Language Models
This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being wide…
Language IdentificationActivity RecognitionDifferentially Private Language Generation and Identification in the Limit
We initiate the study of language generation in the limit, a model recently introduced by Kleinberg and Mullainathan [KM24], under the constraint of differential privacy. We consider the continual release model, where a …
Language IdentificationOn the Price of Privacy for Language Identification and Generation
As large language models (LLMs) are increasingly trained on sensitive user data, understanding the fundamental cost of privacy in language learning becomes essential. We initiate the study of differentially private (DP) …
Language IdentificationBiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement
High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla-English…
Representation LearningLanguage IdentificationText GenerationEnTaCs: Analyzing the Relationship Between Sentiment and Language Choice in English-Tamil Code-Switching
This paper investigates the relationship between utterance sentiment and language choice in English-Tamil code-switched text, using methods from machine learning and statistical modelling. We apply a fine-tuned XLM-RoBER…
Language Identification