paper-with-me

홈 › Papers

Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation

2026-04-27 · Sercan Karakaş, Yusuf Şimşek arxiv

This paper investigates whether source trustworthiness shapes Turkish evidential morphology and whether large language models (LLMs) track this sensitivity. We study the past-domain contrast between -DI and -mIs in controlled cloze contexts where the information source is overtly external, while only its perceived reliability is manipulated (High-Trust vs. Low-Trust). In a human production experiment, native speakers of Turkish show a robust trust effect: High-Trust contexts yield relatively more -DI, whereas Low-Trust contexts yield relatively more -mIs, with the pattern remaining stable across sensitivity analyses. We then evaluate 10 LLMs in three prompting paradigms (open gap-fill, explicit past-tense gap-fill, and forced-choice A/B selection). LLM behavior is highly model- and prompt-dependent: some models show weak or local trust-consistent shifts, but effects are generally unstable, often reversed, and frequently overshadowed by output-compliance problems and strong base-rate suffix preferences. The results provide new evidence for a trust-/commitment-based account of Turkish evidentiality and reveal a clear human-LLM gap in source-sensitive evidential reasoning.

📄 PDF Abstract BibTeX arXiv:2604.24665

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking

2024-05-07 · Emre Can Acikgoz, Mete Erdogan, Deniz Yuret

Large Language Models (LLMs) are becoming crucial across various fields, emphasizing the urgency for high-quality models in underrepresented languages. This study explores the unique challenges faced by low-resource lang…

BenchmarkingModel SelectionTransfer Learning

Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not

2026-04-06 · Sercan Karakaş arxiv

Large language models achieve strong performance on many language tasks, yet it remains unclear whether they integrate world knowledge with syntactic structure in a human-like, structure-sensitive way during ambiguity re…

Mukayese: Turkish NLP Strikes Back

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Having sufficient resources for a language X lifts it from the $\textit{under-resourced}$ languages class, but does not necessarily lift it from the $\textit{under-researched}$ class. In this paper, we address the proble…

BenchmarkingLanguage ModelingLanguage ModellingSentence+1

Mukayese: Turkish NLP Strikes Back

2022-03-02 · Findings (ACL) 2022 5 · Ali Safaya, Emirhan Kurtuluş, Arda Göktoğan, Deniz Yuret

Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the under-researched class. In this paper, we address the problem of the absence of organized benchma…

BenchmarkingLanguage ModelingLanguage ModellingSentence+1

TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish

2024-07-17 · Arda Yüksel, Abdullatif Köksal, Lütfi Kerem Şenel, Anna Korhonen 외

Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs). While existing benchmarks employ automatic translation for multilingual evaluati…

MathMultiple-choiceQuestion Answering