paper-with-me

Papers

Case Study: Testing Model Capabilities in Some Reasoning Tasks

2024-02-15 · Min Zhang, Sato Takumi, Jack Zhang, Jun Wang

Large Language Models (LLMs) excel in generating personalized content and facilitating interactive dialogues, showcasing their remarkable aptitude for a myriad of applications. However, their capabilities in reasoning and providing explainable outputs, especially within the context of reasoning abilities, remain areas for improvement. In this study, we delve into the reasoning abilities of LLMs, highlighting the current challenges and limitations that hinder their effectiveness in complex reasoning scenarios.

📄 PDF Abstract BibTeX arXiv:2402.09967

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LoNLI: An Extensible Framework for Testing Diverse Logical Reasoning Capabilities for NLI

2021-12-04 · Ishan Tarunesh, Somak Aditya, Monojit Choudhury

Natural Language Inference (NLI) is considered a representative task to test natural language understanding (NLU). In this work, we propose an extensible framework to collectively yet categorically test diverse Logical r…

Logical ReasoningNatural Language InferenceNatural Language Understanding

Chain-Of-Thought Prompting Under Streaming Batch: A Case Study

2023-06-01 · Yuxin Tang

Recently, Large Language Models (LLMs) have demonstrated remarkable capabilities. Chain-of-Thought (CoT) has been proposed as a way of assisting LLMs in performing complex reasoning. However, developing effective prompts…

OpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?

2024-11-09 · Leo Li, Ye Luo, Tingyou Pan

The Orion-1 model by OpenAI is claimed to have more robust logical reasoning capabilities than previous large language models. However, some suggest the excellence might be partially due to the model "memorizing" solutio…

Logical ReasoningMath

ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models

2025-05-19 · Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake 외

Chart understanding presents a unique challenge for large vision-language models (LVLMs), as it requires the integration of sophisticated textual and visual reasoning capabilities. However, current LVLMs exhibit a notabl…

Chart Question AnsweringChart UnderstandingQuestion AnsweringVisual Reasoning

ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models

2025-09-29 · Jincheng Liu, Sijun He, Jingjing Wu, Xiangsen Wang 외 arxiv

Recent large language models (LLMs) have shown strong reasoning capabilities. However, a critical question remains: do these models possess genuine strategic reasoning, or do they primarily excel at pattern recognition? …