paper-with-me

홈 › Papers

Prompting with Phonemes: Enhancing LLM Multilinguality for non-Latin Script Languages

2024-11-04 · Hoang Nguyen, Khyati Mahajan, Vikas Yadav, Philip S. Yu, Masoud Hashemi, Rishabh Maheshwary

Multilingual LLMs have achieved remarkable benchmark performance, but we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are pretrained with orthographic scripts, which are dominated by Latin characters that obscure their shared phonology with non-Latin scripts. We propose leveraging phonemic transcriptions as complementary signals to induce script-invariant representations. Our study demonstrates that integrating phonemic signals improves performance across both non-Latin and Latin languages, with a particularly significant impact on closing the performance gap between the two. Through detailed experiments, we show that phonemic and orthographic scripts retrieve distinct examples for in-context learning (ICL). This motivates our proposed Mixed-ICL retrieval strategy, where further aggregation leads to our significant performance improvements for both Latin script languages (up to 12.6%) and non-Latin script languages (up to 15.1%) compared to randomized ICL retrieval.

📄 PDF Abstract BibTeX arXiv:2411.02398

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningRetrieval

Similar Papers 제목 키워드 기반

RomanLens: Latent Romanization and its role in Multilinguality in LLMs

2025-02-11 · Alan Saji, Jaavid Aktar Husain, Thanmay Jayakumar, Raj Dabre 외

Large Language Models (LLMs) exhibit remarkable multilingual generalization despite being predominantly trained on English-centric corpora. A fundamental question arises: how do LLMs achieve such robust multilingual capa…

Language ModelingLanguage Modelling

Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities

2023-12-22 · Yuhao Chen, Chloe Wong, Hanwen Yang, Juan Aguenza 외

This study critically evaluates the efficacy of prompting methods in enhancing the mathematical reasoning capability of large language models (LLMs). The investigation uses three prescriptive prompting methods - simple, …

ChatbotGSM8KLanguage ModellingMath+3

Towards Zero-shot Learning for Automatic Phonemic Transcription

2020-02-26 · Xinjian Li, Siddharth Dalmia, David R. Mortensen, Juncheng Li 외

Automatic phonemic transcription tools are useful for low-resource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, mult…

Zero-Shot Learning

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet

2026-06-18 · Milan Miletić, Julie Kallini, Ekaterina Shutova arxiv

Multilingual language models often exhibit performance disparities across languages that can arise as early as the tokenization stage. Widely-used subword tokenization approaches favor high-resource languages, and tokeni…

Stochastic model for phonemes uncovers an author-dependency of their usage

2015-10-05 · Weibing Deng, Armen E. Allahverdyan

We study rank-frequency relations for phonemes, the minimal units that still relate to linguistic meaning. We show that these relations can be described by the Dirichlet distribution, a direct analogue of the ideal-gas m…