paper-with-me

홈 › Papers

LLMs Know More About Numbers than They Can Say

2026-02-08 · Fengting Yuchi, Li Du, Jason Eisner arxiv

Although state-of-the-art LLMs can solve math problems, we find that they make errors on numerical comparisons with mixed notation: "Which is larger, $5.7 \times 10^2$ or $580$?" This raises a fundamental question: Do LLMs even know how big these numbers are? We probe the hidden states of several smaller open-source LLMs. A single linear projection of an appropriate hidden layer encodes the log-magnitudes of both kinds of numerals, allowing us to recover the numbers with relative error of about 2.3% (on restricted synthetic text) or 19.06% (on scientific papers). Furthermore, the hidden state after reading a pair of numerals encodes their ranking, with a linear classifier achieving over 90% accuracy. Yet surprisingly, when explicitly asked to rank the same pairs of numerals, these LLMs achieve only 50-70% accuracy, with worse performance for models whose probes are less effective. Finally, we show that incorporating the classifier probe's log-loss as an auxiliary objective during finetuning brings an additional 3.22% improvement in verbalized accuracy over base models, demonstrating that improving models' internal magnitude representations can enhance their numerical reasoning capabilities. Our code is available at https://github.com/VCY019/Numeracy-Probing.

📄 PDF Abstract BibTeX arXiv:2602.07812

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles

2025-06-16 · Antara Raaghavi Bhattacharya, Isabel Papadimitriou, Kathryn Davidson, David Alvarez-Melis

Across languages, numeral systems vary widely in how they construct and combine numbers. While humans consistently learn to navigate this diversity, large language models (LLMs) struggle with linguistic-mathematical puzz…

DiversityMathematical ReasoningNavigate

Language Models Encode Numbers Using Digit Representations in Base 10

2024-10-15 · Amit Arnold Levy, Mor Geva

Large language models (LLMs) frequently make errors when handling even simple numerical problems, such as comparing two small numbers. A natural hypothesis is that these errors stem from how LLMs represent numbers, and s…

Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

2025-05-28 · Michael Kirchhof, Gjergji Kasneci, Enkelejda Kasneci

Large-language models (LLMs) and chatbot agents are known to provide wrong outputs at times, and it was recently found that this can never be fully prevented. Hence, uncertainty quantification plays a crucial role, aimin…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model+2

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents

2026-01-07 · Alessio Buscemi, Daniele Proverbio, Alessandro Di Stefano, The-Anh Han 외 arxiv

LLMs-based agents increasingly operate in multi-agent environments where strategic interaction and coordination are required. While existing work has largely focused on individual agents or on interacting agents sharing …

Bounding the number of reticulation events for displaying multiple trees in a phylogenetic network

2024-08-26 · Yufeng Wu, Louxin Zhang

Reconstructing a parsimonious phylogenetic network that displays multiple phylogenetic trees is an important problem in theory of phylogenetics, where the complexity of the inferred networks is measured by reticulation n…