paper-with-me

홈 › Papers

Do LLMs Recognize me, When I is not me: Assessment of LLMs Understanding of Turkish Indexical Pronouns in Indexical Shift Contexts

2024-06-08 · Metehan Oğuz, Yusuf Umut Ciftci, Yavuz Faruk Bakman

Large language models (LLMs) have shown impressive capabilities in tasks such as machine translation, text summarization, question answering, and solving complex mathematical problems. However, their primary training on data-rich languages like English limits their performance in low-resource languages. This study addresses this gap by focusing on the Indexical Shift problem in Turkish. The Indexical Shift problem involves resolving pronouns in indexical shift contexts, a grammatical challenge not present in high-resource languages like English. We present the first study examining indexical shift in any language, releasing a Turkish dataset specifically designed for this purpose. Our Indexical Shift Dataset consists of 156 multiple-choice questions, each annotated with necessary linguistic details, to evaluate LLMs in a few-shot setting. We evaluate recent multilingual LLMs, including GPT-4, GPT-3.5, Cohere-AYA, Trendyol-LLM, and Turkcell-LLM, using this dataset. Our analysis reveals that even advanced models like GPT-4 struggle with the grammatical nuances of indexical shift in Turkish, achieving only moderate performance. These findings underscore the need for focused research on the grammatical challenges posed by low-resource languages. We released the dataset and code \href{https://anonymous.4open.science/r/indexical_shift_llm-E1B4} {here}.

📄 PDF Abstract BibTeX arXiv:2406.05569

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationMultiple-choiceQuestion AnsweringText Summarization

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Weight Decay 설명 없음
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…

Similar Papers 제목 키워드 기반

The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding

2025-02-13 · Mo Yu, Lemao Liu, Junjie Wu, Tsz Ting Chung 외

In a systematic way, we investigate a widely asked question: Do LLMs really understand what they say?, which relates to the more familiar term Stochastic Parrot. To this end, we propose a summative assessment over a care…

In-Context LearningMemorization

PRISON: Unmasking the Criminal Potential of Large Language Models

2025-06-19 · Xinyi Wu, Geng Hong, Pei Chen, Yueyue Chen 외

As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked the systematic understanding and assessment of their criminal capability in real…

Adversarial Robustness

Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models

2024-05-25 · Abhishek Kumar, Robert Morabito, Sanzhar Umbet, Jad Kabbara 외

As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the outp…

Risk Assessment Framework for Code LLMs via Leveraging Internal States

2025-04-20 · Yuheng Huang, Lei Ma, Keizaburo Nishikino, Takumi Akazaki

The pre-training paradigm plays a key role in the success of Large Language Models (LLMs), which have been recognized as one of the most significant advancements of AI recently. Building on these breakthroughs, code LLMs…

Unsupervised Pre-training

FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation

2026-02-25 · Xiyuan Zhang, Huihang Wu, Jiayu Guo, Zhenlin Zhang 외 arxiv

We introduce FIRE, a comprehensive benchmark designed to evaluate both the theoretical financial knowledge of LLMs and their ability to handle practical business scenarios. For theoretical assessment, we curate a diverse…