paper-with-me

Papers

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

2025-07-16 · Ziyu Ge, Gabriel Chua, Leanne Tan, Roy Ka-Wei Lee arxiv

As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing, and culturally embedded markers of harmful speech. Translating toxic content between low-resource language pairs poses additional challenges due to scarce parallel data and safety filters that sanitize offensive expressions. In this work, we propose a reproducible, two-stage framework for toxicity-preserving translation, demonstrated on a code-mixed Singlish safety corpus. First, we perform human-verified few-shot prompt engineering: we iteratively curate and rank annotator-selected Singlish-target examples to capture nuanced slang, tone, and toxicity. Second, we optimize model-prompt pairs by benchmarking several large language models using semantic similarity via direct and back-translation. Quantitative human evaluation confirms the effectiveness and efficiency of our pipeline. Beyond improving translation quality, our framework contributes to the safety of multicultural LLMs by supporting culturally sensitive moderation and benchmarking in low-resource contexts. By positioning Singlish as a testbed for inclusive NLP, we underscore the importance of preserving sociolinguistic nuance in real-world applications such as content moderation and regional platform governance.

📄 PDF Abstract BibTeX arXiv:2507.11966

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilarityPrompt Engineering

Similar Papers 제목 키워드 기반

Singlish Where Got Rules One? Constructing a Computational Grammar for Singlish

2022-06-01 · LREC 2022 6 · Siew Yeng Chow, Francis Bond

Singlish is a variety of English spoken in Singapore. In this paper, we share some of its grammar features and how they are implemented in the construction of a computational grammar of Singlish as a branch of English gr…

RabakBench: Scaling Human Annotations to Construct Localized Multilingual Safety Benchmarks for Low-Resource Languages

2025-07-08 · Gabriel Chua, Leanne Tan, Ziyu Ge, Roy Ka-Wei Lee

Large language models (LLMs) and their safety classifiers often perform poorly on low-resource languages due to limited training data and evaluation benchmarks. This paper introduces RabakBench, a new multilingual safety…

Red Teaming

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

2025-09-18 · Yujia Hu, Ming Shan Hee, Preslav Nakov, Roy Ka-Wei Lee arxiv

The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap.…

What talking you?: Translating Code-Mixed Messaging Texts to English

2024-11-08 · Lynnette Hui Xian Ng, Luo Qi Chan

Translation of code-mixed texts to formal English allow a wider audience to understand these code-mixed languages, and facilitate downstream analysis applications such as sentiment analysis. In this work, we look at tran…

Sentiment AnalysisTranslation

Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling

2026-01-05 · Berk Atil, Rebecca J. Passonneau, Ninareh Mehrabi arxiv

Toxicity detection is inherently subjective, shaped by the diverse perspectives and social priors of different demographic groups. While ``pluralistic'' modeling as used in economics and the social sciences aims to captu…