paper-with-me

Papers

ChiEngMixBench: Evaluating Large Language Models on Spontaneous and Natural Chinese-English Code-Mixed Generation

2026-01-02 · Qingyan Yang, Tongxi Wang, Yunsheng Luo arxiv

Code-mixing is increasingly prevalent in interactions between humans and large language models, yet existing work often reduces it to a translation or convertibility problem, making it difficult to assess whether a model's switching behavior is context-appropriate and aligned with human conventions. We introduce ChiEngMixBench, the first benchmark designed to evaluate code-mixing ability in authentic community contexts, built upon a general construction pipeline that enables scalable dataset development across domains and bilingual pairs. ChiEngMixBench formulates code-mixing as a cognitive alignment problem, characterized by two complementary signals: Spontaneity and Naturalness. Empirical evaluation shows that our metrics can systematically distinguish code-mixing performance across models. Beyond benchmarking, we further uncover an implicitly emergent Terminology Layering Strategy, a phenomenon consistent with the Matrix Language Frame (MLF) theory, indicating structured cognitive alignment between multilingual large language models and human communication.

📄 PDF Abstract BibTeX arXiv:2601.16217

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

2024-07-18 · Weiqin Li, Peiji Yang, Yicheng Zhong, Yixuan Zhou 외

Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS sy…

Language ModelingLanguage ModellingSpeech Synthesistext-to-speech+2

Evaluating Sampling-based Filler Insertion with Spontaneous TTS

2022-06-01 · LREC 2022 6 · Siyang Wang, Joakim Gustafson, Éva Székely

Inserting fillers (such as “um”, “like”) to clean speech text has a rich history of study. One major application is to make dialogue systems sound more spontaneous. The ambiguity of filler occurrence and inter-speaker di…

Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility

2025-05-22 · Sheng-Fu Wang, Laurent Prevot, Jou-an Chi, Ri-Sheng Huang 외

The achievements of Large Language Models in Natural Language Processing, especially for high-resource languages, call for a better understanding of their characteristics from a cognitive perspective. Researchers have at…

Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model

2024-05-16 · Siyang Wang, Éva Székely

Recent advances in generative language modeling applied to discrete speech tokens presented a new avenue for text-to-speech (TTS) synthesis. These speech language models (SLMs), similarly to their textual counterparts, a…

HallucinationLanguage ModelingLanguage ModellingSpeech Synthesis+3

CASPER: A Large Scale Spontaneous Speech Dataset

2025-05-30 · Cihan Xiao, Ruixing Liang, Xiangyu Zhang, Mehmet Emre Tiryaki 외

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets c…