paper-with-me

Papers

Evaluating LLMs on Chinese Topic Constructions: A Research Proposal Inspired by Tian et al. (2024)

2025-04-21 · Xiaodong Yang

This paper proposes a framework for evaluating large language models (LLMs) on Chinese topic constructions, focusing on their sensitivity to island constraints. Drawing inspiration from Tian et al. (2024), we outline an experimental design for testing LLMs' grammatical knowledge of Mandarin syntax. While no experiments have been conducted yet, this proposal aims to provide a foundation for future studies and invites feedback on the methodology.

📄 PDF Abstract BibTeX arXiv:2504.14969

Code (0)

등록된 구현이 없습니다.

Tasks

Experimental DesignSensitivity

Similar Papers 제목 키워드 기반

Topic-comment constructions in L1-Chinese learners’ English

2018-12-01 · PACLIC 2018 12 · Yicheng Rong, Hui Chang, Candice Chi-Hang Cheung

CHBench: A Chinese Dataset for Evaluating Health in Large Language Models

2024-09-24 · Chenlu Guo, Nuo Xu, Yi Chang, Yuan Wu

With the rapid development of large language models (LLMs), assessing their performance on health-related inquiries has become increasingly essential. It is critical that these models provide accurate and trustworthy hea…

Misinformation

CARE-MI: Chinese Benchmark for Misinformation Evaluation in Maternity and Infant Care

2023-07-04 · NeurIPS 2023 11 · Tong Xiang, Liangzhi Li, Wangyue Li, Mingbai Bai 외

The recent advances in natural language processing (NLP), have led to a new trend of applying large language models (LLMs) to real-world scenarios. While the latest LLMs are astonishingly fluent when interacting with hum…

Misinformation

Evaluating Chinese Ambiguity Understanding in Large Language Models

2026-05-15 · Junwen Mo, Yuanzhi Lu, Yifang Xue, Ke Xu 외 arxiv

Linguistic ambiguity is critical to the robustness of Large Language Models (LLMs), yet existing research focuses mostly on English, with limited attention devoted to Chinese. Existing Chinese ambiguity datasets (e.g., C…

Machine Translation

ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models

2024-10-24 · Hengxiang Zhang, Hongfu Gao, Qiang Hu, Guanhua Chen 외

With the rapid development of Large language models (LLMs), understanding the capabilities of LLMs in identifying unsafe content has become increasingly important. While previous works have introduced several benchmarks …