paper-with-me

Papers

LLMs Struggle with NLI for Perfect Aspect: A Cross-Linguistic Study in Chinese and Japanese

2025-08-16 · Jie Lu, Du Jin, Hitomi Yanaka arxiv

Unlike English, which uses distinct forms (e.g., had, has, will have) to mark the perfect aspect across tenses, Chinese and Japanese lack separate grammatical forms for tense within the perfect aspect, which complicates Natural Language Inference (NLI). Focusing on the perfect aspect in these languages, we construct a linguistically motivated, template-based NLI dataset (1,350 pairs per language). Experiments reveal that even advanced LLMs struggle with temporal inference, particularly in detecting subtle tense and reference-time shifts. These findings highlight model limitations and underscore the need for cross-linguistic evaluation in temporal semantics. Our dataset is available at https://github.com/Lujie2001/CrossNLI.

📄 PDF Abstract BibTeX arXiv:2508.11927

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Inference

Similar Papers 제목 키워드 기반

How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs

2025-07-18 · Karin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu 외 arxiv

Large language models (LLMs) exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these behaviors reflect human-like cognition versus advanced pattern recognition remains an open question. …

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

2025-11-04 · Lachlan McPheat, Navdeep Kaur, Robert Blackwell, Alessandra Russo 외 arxiv

We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning ability. The generation of DecompSR allows …

Spatial Reasoning

Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting

2025-10-23 · Josh McGiff, Khanh-Tung Tran, William Mulcahy, Dáibhidh Ó Luinín 외 arxiv

We present Irish-BLiMP (Irish Benchmark of Linguistic Minimal Pairs), the first dataset and framework designed for fine-grained evaluation of linguistic competence in the Irish language, an endangered language. Drawing o…

Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles

2025-06-16 · Antara Raaghavi Bhattacharya, Isabel Papadimitriou, Kathryn Davidson, David Alvarez-Melis

Across languages, numeral systems vary widely in how they construct and combine numbers. While humans consistently learn to navigate this diversity, large language models (LLMs) struggle with linguistic-mathematical puzz…

DiversityMathematical ReasoningNavigate

Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability

2024-09-24 · Xufeng Duan, Xinyu Zhou, Bei Xiao, Zhenguang G. Cai

As large language models (LLMs) advance in their linguistic capacity, understanding how they capture aspects of language competence remains a significant challenge. This study therefore employs psycholinguistic paradigms…

Language ModelingLanguage Modelling