paper-with-me

Papers

LTD-Bench: Evaluating Large Language Models by Letting Them Draw

2025-11-04 · Liuhao Lin, Ke Li, Zihan Xu, Yuchen Shi, Yulei Qin, Yan Zhang, Xing Sun, Rongrong Ji arxiv

Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research--relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive understanding of model capabilities. This deficiency creates a dangerous disconnect between reported performance and practical abilities, particularly for applications requiring physical world understanding. We introduce LTD-Bench, a breakthrough benchmark that transforms LLM evaluation from abstract scores to directly observable visual outputs by requiring models to generate drawings through dot matrices or executable code. This approach makes spatial reasoning limitations immediately apparent even to non-experts, bridging the fundamental gap between statistical performance and intuitive assessment. LTD-Bench implements a comprehensive methodology with complementary generation tasks (testing spatial imagination) and recognition tasks (assessing spatial perception) across three progressively challenging difficulty levels, methodically evaluating both directions of the critical language-spatial mapping. Our extensive experiments with state-of-the-art models expose an alarming capability gap: even LLMs achieving impressive results on traditional benchmarks demonstrate profound deficiencies in establishing bidirectional mappings between language and spatial concept--a fundamental limitation that undermines their potential as genuine world models. Furthermore, LTD-Bench's visual outputs enable powerful diagnostic analysis, offering a potential approach to investigate model similarity.

📄 PDF Abstract BibTeX arXiv:2511.02347

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

ChainStream: An LLM-based Framework for Unified Synthetic Sensing

2024-12-13 · Jiacheng Liu, Yuanchun Li, Liangyan Li, Yi Sun 외

Many applications demand context sensing to offer personalized and timely services. Yet, developing sensing programs can be challenging for developers and using them is privacy-concerning for end-users. In this paper, we…

Code Generation

Earth Embeddings

2026-08-04 · Adam J. Stewart, Heng Fang, Isaac A. Corley, Xiao Xiang Zhu arxiv

Earth observation is moving from foundation models that users must run themselves toward embedding products that package model feature outputs as reusable data without needing to download and process the imagery used to …

LettinGo: Explore User Profile Generation for Recommendation System

2025-06-23 · Lu Wang, Di Zhang, Fangkai Yang, Pu Zhao 외

User profiling is pivotal for recommendation systems, as it transforms raw user interaction data into concise and structured representations that drive personalized recommendations. While traditional embedding-based prof…

Profile GenerationRecommendation Systems

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

2026-04-20 · Shaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton 외 arxiv

Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a h…

Mathematical Reasoning

RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics

2025-05-18 · Jie Zhang, Cezara Petrui, Kristina Nikolić, Florian Tramèr

Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions -- failing to capture the nature of m…

Mathematical Reasoning