paper-with-me

Papers

Improving Language Identification for Code-Switched Utterances with Integer Linear Programming

2026-09-04 · Joanna Radoła, Josep Maria Crego, François Yvon arxiv

Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this paper, we revisit MaskLID, a state-of-the art approach for CS identification, which requires no training and detects arbitrary language combinations. We make three main contributions: (a) we reveal, and address, a major issue of MaskLID: its overreliance on word-level language association scores; (b) we reformulate the underlying optimization algorithm as an Integer Linear Program, enabling us to experiment with a large set of clear and interpretable constraints; (c) each of these improvements vastly improves the baseline system, as we illustrate in experiments involving 10~diverse languages, where we observe a strong boost in performance on CS benchmarks. We release our code and data for reproducibility.

📄 PDF Abstract BibTeX arXiv:2609.05099

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

MERLIon CCS Challenge Evaluation Plan

2023-05-31 · Leibny Paola Garcia Perera, Y. H. Victoria Chua, Hexin Liu, Fei Ting Woon 외

This paper introduces the inaugural Multilingual Everyday Recordings- Language Identification on Code-Switched Child-Directed Speech (MERLIon CCS) Challenge, focused on developing robust language identification and langu…

Language IdentificationTask 2

CST5: Data Augmentation for Code-Switched Semantic Parsing

2022-11-14 · Anmol Agarwal, Jigar Gupta, Rahul Goel, Shyam Upadhyay 외

Extending semantic parsers to code-switched input has been a challenging problem, primarily due to a lack of supervised training data. In this work, we introduce CST5, a new data augmentation technique that finetunes a T…

Data AugmentationSemantic Parsing

CST5: Data augmentation for Code-Switched Semantic Parsing

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Extending semantic parsers to code-switched input has been a challenging problem, primarily due to lack of labeled data for supervision. In this work, we introduce CST5, a new data augmentation technique that finetunes …

Data AugmentationSemantic Parsing

EnTaCs: Analyzing the Relationship Between Sentiment and Language Choice in English-Tamil Code-Switching

2026-03-27 · Paul Bontempo arxiv

This paper investigates the relationship between utterance sentiment and language choice in English-Tamil code-switched text, using methods from machine learning and statistical modelling. We apply a fine-tuned XLM-RoBER…

Language Identification

Simple Features for Strong Performance on Named Entity Recognition in Code-Switched Twitter Data

2018-07-01 · WS 2018 7 · Devanshu Jain, Maria Kustikova, Mayank Darbari, Rishabh Gupta 외

In this work, we address the problem of Named Entity Recognition (NER) in code-switched tweets as a part of the Workshop on Computational Approaches to Linguistic Code-switching (CALCS) at ACL{'}18. Code-switching is the…

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4