paper-with-me

홈 › Papers

Dialectical language model evaluation: An initial appraisal of the commonsense spatial reasoning abilities of LLMs

2023-04-22 · Anthony G Cohn, Jose Hernandez-Orallo

Language models have become very popular recently and many claims have been made about their abilities, including for commonsense reasoning. Given the increasingly better results of current language models on previous static benchmarks for commonsense reasoning, we explore an alternative dialectical evaluation. The goal of this kind of evaluation is not to obtain an aggregate performance value but to find failures and map the boundaries of the system. Dialoguing with the system gives the opportunity to check for consistency and get more reassurance of these boundaries beyond anecdotal evidence. In this paper we conduct some qualitative investigations of this kind of evaluation for the particular case of spatial reasoning (which is a fundamental aspect of commonsense reasoning). We conclude with some suggestions for future work both to improve the capabilities of language models and to systematise this kind of dialectical evaluation.

📄 PDF Abstract BibTeX arXiv:2304.11164

Code (0)

등록된 구현이 없습니다.

Tasks

Language Model EvaluationLanguage ModelingLanguage ModellingSpatial Reasoning

Similar Papers 제목 키워드 기반

A Semi-supervised Approach for a Better Translation of Sentiment in Dialectical Arabic UGT

2022-10-21 · Hadeel Saadany, Constantin Orasan, Emad Mohamed, Ashraf Tantawy

In the online world, Machine Translation (MT) systems are extensively used to translate User-Generated Text (UGT) such as reviews, tweets, and social media posts, where the main message is often the author's positive or …

Language ModellingMachine TranslationNMTTranslation

Multi-Agent Dialectical Refinement for Enhanced Argument Classification

2026-03-29 · Jakub Bąba, Jarosław A. Chudziak arxiv

Argument Mining (AM) is a foundational technology for automated writing evaluation, yet traditional supervised approaches rely heavily on expensive, domain-specific fine-tuning. While Large Language Models (LLMs) offer a…

Component ClassificationArgument Mining

Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring

2026-03-28 · Jakub Masłowski, Jarosław A. Chudziak arxiv

Large Language Models (LLMs) are being increasingly used as autonomous agents in complex reasoning tasks, opening the niche for dialectical interactions. However, Multi-Agent systems implemented with systematically uncon…

Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided

2024-04-01 · Hongli Zhan, Allen Zheng, Yoon Kyung Lee, Jina Suh 외

Large language models (LLMs) have offered new opportunities for emotional support, and recent work has shown that they can produce empathic responses to people in distress. However, long-term mental well-being requires e…

DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models

2025-01-27 · Niyati Bafna, Emily Chang, Nathaniel R. Robinson, David R. Mortensen 외

Most of the world's languages and dialects are low-resource, and lack support in mainstream machine translation (MT) models. However, many of them have a closely-related high-resource language (HRL) neighbor, and differ …

Machine Translation