paper-with-me

홈 › Papers

Too long; didn't solve

2026-04-08 · Lucía M. Cabrera, Isaac Saxton-Knight, Jocelyn D'Arcy arxiv

Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language models, yet little is known about how their structural properties influence model behaviour. In this work, we investigate two structural length variables, prompt length and solution length, and analyse how they relate to model performance on a newly constructed adversarial dataset of expert-authored mathematics problems. We find that both prompt and solution lengths correlate positively with increased model failure across models. We also include a secondary, exploratory analysis of cross-model disagreement. Under a difficulty-adjusted normalised analysis, both variables retain weak negative associations with realised model separation, slightly stronger for prompt length. Overall, our main robust finding is that structural length is linked to empirical difficulty in this dataset.

📄 PDF Abstract BibTeX arXiv:2604.07593

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CDIDN: A Registration Model with High Deformation Impedance Capability for Long-Term Tracking of Pulmonary Lesion Dynamics

2023-05-18 · Xinyu Zhao, Sa Huang, Wei Pang, You Zhou

We study the problem of registration for medical CT images from a novel perspective -- the sensitivity to degree of deformations in CT images. Although some learning-based methods have shown success in terms of average a…

Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels

2025-05-20 · Sil Hamilton, Rebecca M. M. Hicke, Matthew Wilkens, David Mimno

Although the context length of large language models (LLMs) has increased to millions of tokens, evaluating their effectiveness beyond needle-in-a-haystack approaches has proven difficult. We argue that novels provide a …

Language ModelingLanguage ModellingLong-Context Understanding

Co-Attention Hierarchical Network: Generating Coherent Long Distractors for Reading Comprehension

2019-11-20 · Xiaorui Zhou, Senlin Luo, Yunfang Wu

In reading comprehension, generating sentence-level distractors is a significant task, which requires a deep understanding of the article and question. The traditional entity-centered methods can only generate word-level…

DecoderDistractor GenerationReading ComprehensionSemantic Similarity+2

Modeling High-order Interactions across Multi-interests for Micro-video Recommendation

2021-04-01 · Dong Yao, Shengyu Zhang, Zhou Zhao, Wenyan Fan 외

Personalized recommendation system has become pervasive in various video platform. Many effective methods have been proposed, but most of them didn't capture the user's multi-level interest trait and dependencies between…

Vocal Bursts Intensity Prediction

Task Proposal: The TL;DR Challenge

2018-11-01 · WS 2018 11 · Shahbaz Syed, Michael V{\"o}lske, Martin Potthast, Nedim Lipka 외

The TL;DR challenge fosters research in abstractive summarization of informal text, the largest and fastest-growing source of textual data on the web, which has been overlooked by summarization research so far. The chall…

Abstractive Text SummarizationInformation RetrievalText GenerationText Summarization