paper-with-me

Papers Language Identification

“Language Identification” 태그가 달린 논문 851편 · 필터 해제

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

2026-09-07 · Gonçalo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos 외 arxiv

Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline t…

Language Identification

Improving Language Identification for Code-Switched Utterances with Integer Linear Programming

2026-09-04 · Joanna Radoła, Josep Maria Crego, François Yvon arxiv

Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this pa…

Language Identification

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

2026-08-11 · Chris Han, Pengzhi Gao, Pei Fu, Jian Luan hf

We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO)…

Language IdentificationReinforcement LearningMachine Translation

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

2026-07-16 · Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen arxiv

LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, l…

Language Identification

Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry

2026-07-16 · Paul-Andrei Pogăcean, Sanda-Maria Avram arxiv

Language identification is commonly addressed using either neural architectures or statistical n-gram models. Neural approaches typically require substantial computational resources, whereas classical frequency-based met…

Language Identification

Language Identification with Succinct Machine-Independent Traces

2026-07-14 · Moses Charikar, Jon Kleinberg, Chirag Pabbaraju arxiv

Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negativ…

Language Identification

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages

2026-07-01 · Jesujoba O. Alabi, Julian Herreilers, Badr M. Abdullah, Dietrich Klakow arxiv

Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architec…

Language IdentificationSpeech Recognition

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

2026-06-30 · Nina Hosseini-Kivanani, Sandipana Dowerah arxiv

State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, …

Language IdentificationCross-Lingual TransferActivity Detection

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

2026-06-16 · Reihaneh Amooie, Yun Hao, Wietse de Vries, Jelske Dijkstra 외 arxiv

This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a…

Language IdentificationSpeech Recognition

Generating in the Limit with Infinitely Many Hallucinations

2026-06-08 · Irene Strauss, Alexandra Butoi, Ryan Cotterell arxiv

The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language. The r…

Language Identification

CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios

2026-06-04 · Michal Tichý, Jindřich Libovický arxiv

We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language identification: cousin languages and orthographic noise. Our dataset has…

Language Identification

The Inclusion Depth of Pattern Languages: An Open Problem in Algorithmic Learning Theory

2026-05-28 · Wei Luo arxiv

Pattern languages are a classical model in formal language theory and algorithmic learning theory. This note formulates the problem of computing the inclusion depth of a pattern language: the length of the longest strict…

Language Identification

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

2026-05-26 · Hanif Rahman arxiv

Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A system may produce no audio, speak a neighbouring language, preserve…

Language Identification

Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus

2026-05-21 · Wajdi Zaghouani, Mabrouka Bessghaier, MD. Rafiul Biswas, Shimaa Amer Ibrahim arxiv

This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corpus was collected from 51,660 pages acros…

Language Identification

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

2026-05-21 · Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel, Josef van Genabith 외 arxiv

Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unreliable in multilingual settings: most SAEs …

Language IdentificationMachine Translation

Demographic and Linguistic Bias Evaluation in Omnimodal Language Models

2026-04-11 · Alaa Elobaid arxiv

This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being wide…

Language IdentificationActivity Recognition

Differentially Private Language Generation and Identification in the Limit

2026-04-09 · Anay Mehrotra, Grigoris Velegkas, Xifan Yu, Felix Zhou arxiv

We initiate the study of language generation in the limit, a model recently introduced by Kleinberg and Mullainathan [KM24], under the constraint of differential privacy. We consider the continual release model, where a …

Language Identification

On the Price of Privacy for Language Identification and Generation

2026-04-08 · Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao arxiv

As large language models (LLMs) are increasingly trained on sensitive user data, understanding the fundamental cost of privacy in language learning becomes essential. We initiate the study of differentially private (DP) …

Language Identification

BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement

2026-04-06 · Abdullah Al Shafi, Swapnil Kundu Argha, M. A. Moyeen, Abdul Muntakim 외 arxiv

High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla-English…

Representation LearningLanguage IdentificationText Generation

EnTaCs: Analyzing the Relationship Between Sentiment and Language Choice in English-Tamil Code-Switching

2026-03-27 · Paul Bontempo arxiv

This paper investigates the relationship between utterance sentiment and language choice in English-Tamil code-switched text, using methods from machine learning and statistical modelling. We apply a fine-tuned XLM-RoBER…

Language Identification
1–20 / 851 다음 →