paper-with-me

Language Identification

6개 벤치마크 · 논문 851편 · 이 태스크의 논문 보기 →

Benchmarks

VOXLINGUA107

결과 2개

GlotLID-C

결과 1개

OpenSubtitles

결과 1개

VoxForge

결과 1개

Most implemented

Papers

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

2026-09-07 · Gonçalo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos 외 arxiv

Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline t…

Language Identification

Improving Language Identification for Code-Switched Utterances with Integer Linear Programming

2026-09-04 · Joanna Radoła, Josep Maria Crego, François Yvon arxiv

Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this pa…

Language Identification

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

2026-08-11 · Chris Han, Pengzhi Gao, Pei Fu, Jian Luan hf

We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO)…

Language IdentificationReinforcement LearningMachine Translation

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

2026-07-16 · Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen arxiv

LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, l…

Language Identification

Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry

2026-07-16 · Paul-Andrei Pogăcean, Sanda-Maria Avram arxiv

Language identification is commonly addressed using either neural architectures or statistical n-gram models. Neural approaches typically require substantial computational resources, whereas classical frequency-based met…

Language Identification

Language Identification with Succinct Machine-Independent Traces

2026-07-14 · Moses Charikar, Jon Kleinberg, Chirag Pabbaraju arxiv

Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negativ…

Language Identification

전체 851편 보기 →