paper-with-me

홈 › Papers

Negation: A Pink Elephant in the Large Language Models' Room?

2025-03-28 · Tereza Vrabcová, Marek Kadlčík, Petr Sojka, Michal Štefánik, Michal Spiegel

Negations are key to determining sentence meaning, making them essential for logical reasoning. Despite their importance, negations pose a substantial challenge for large language models (LLMs) and remain underexplored. We construct two multilingual natural language inference (NLI) datasets with \textit{paired} examples differing in negation. We investigate how model size and language impact its ability to handle negation correctly by evaluating popular LLMs. Contrary to previous work, we show that increasing the model size consistently improves the models' ability to handle negations. Furthermore, we find that both the models' reasoning accuracy and robustness to negation are language-dependent and that the length and explicitness of the premise have a greater impact on robustness than language. Our datasets can facilitate further research and improvements of language model reasoning in multilingual settings.

📄 PDF Abstract BibTeX arXiv:2503.22395

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLogical ReasoningNatural Language InferenceNegationSentence

Similar Papers 제목 키워드 기반

Suppressing Pink Elephants with Direct Principle Feedback

2024-02-12 · Louis Castricato, Nathan Lile, Suraj Anand, Hailey Schoelkopf 외

Existing methods for controlling language models, such as RLHF and Constitutional AI, involve determining which LLM behaviors are desirable and training them into a language model. However, in many cases, it is desirable…

Language ModelingLanguage Modelling

Do not think about pink elephant!

2024-04-22 · Kyomin Hwang, Suyoung Kim, JunHoo Lee, Nojun Kwak

Large Models (LMs) have heightened expectations for the potential of general AI as they are akin to human intelligence. This paper shows that recent large models such as Stable Diffusion and DALL-E3 also share the vulner…

Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment

2024-09-26 · Yan Liu, Xiaoyuan Yi, Xiaokang Chen, Jing Yao 외

The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models for optimization or evaluation, neglect…

Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room

2025-04-22 · Danial Hooshyar, Gustav Šír, Yeongwook Yang, Eve Kikas 외

Despite significant advancements in AI-driven educational systems and ongoing calls for responsible AI for education, several critical issues remain unresolved -- acting as the elephant in the room within AI in education…

BenchmarkingFairness

Can We Catch the Elephant? A Survey of the Evolvement of Hallucination Evaluation on Natural Language Generation

2024-04-18 · Siya Qi, Yulan He, Zheng Yuan

Hallucination in Natural Language Generation (NLG) is like the elephant in the room, obvious but often overlooked until recent achievements significantly improved the fluency and grammaticality of generated text. As the …

HallucinationHallucination EvaluationSurveyText Generation