paper-with-me

Papers

Human Behavioral Benchmarking: Numeric Magnitude Comparison Effects in Large Language Models

2023-05-18 · Raj Sanjay Shah, Vijay Marupudi, Reba Koenen, Khushi Bhardwaj, Sashank Varma

Large Language Models (LLMs) do not differentially represent numbers, which are pervasive in text. In contrast, neuroscience research has identified distinct neural representations for numbers and words. In this work, we investigate how well popular LLMs capture the magnitudes of numbers (e.g., that $4 < 5$) from a behavioral lens. Prior research on the representational capabilities of LLMs evaluates whether they show human-level performance, for instance, high overall accuracy on standard benchmarks. Here, we ask a different question, one inspired by cognitive science: How closely do the number representations of LLMscorrespond to those of human language users, who typically demonstrate the distance, size, and ratio effects? We depend on a linking hypothesis to map the similarities among the model embeddings of number words and digits to human response times. The results reveal surprisingly human-like representations across language models of different architectures, despite the absence of the neural circuitry that directly supports these representations in the human brain. This research shows the utility of understanding LLMs using behavioral benchmarks and points the way to future work on the number representations of LLMs and their cognitive plausibility.

📄 PDF Abstract BibTeX arXiv:2305.10782

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Number Representations in LLMs: A Computational Parallel to Human Perception

2025-02-22 · H. V. AlquBoj, Hilal AlQuabeh, Velibor Bojkovic, Tatsuya Hiraoka 외

Humans are believed to perceive numbers on a logarithmic mental number line, where smaller values are represented with greater resolution than larger ones. This cognitive bias, supported by neuroscience and behavioral st…

Dimensionality Reduction

Metrics vs Surveys: An Analysis for Human-Aligned Benchmarking in Social Robot Navigation

2025-10-03 · Stefano Trepella, Mauro Martini, Noé Pérez-Higueras, Andrea Ostuni 외 arxiv

Social, also called human-aware, navigation is a key challenge for integrating mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be …

Robot Navigation

Towards a Multidimensional Evaluation Framework for Empathetic Conversational Systems

2024-07-26 · Aravind Sesagiri Raamkumar, Siyuan Brandon Loh

Empathetic Conversational Systems (ECS) are built to respond empathetically to the user's emotions and sentiments, regardless of the application domain. Current ECS studies evaluation approaches are restricted to offline…

Benchmarking

Perception of visual numerosity in humans and machines

2019-07-16 · Alberto Testolin, Serena Dolfi, Mathijs Rochus, Marco Zorzi

Numerosity perception is foundational to mathematical learning, but its computational bases are strongly debated. Some investigators argue that humans are endowed with a specialized system supporting numerical representa…

Flaws in the LLM Automation Narrative

2026-06-09 · George Perrett, Javae Elliott, Jennifer Hill, Marc Scott arxiv

Large Language Models (LLMs) are increasingly described as performing at the level of human experts on knowledge economy tasks. These claims are primarily based on how LLMs perform on benchmarking tasks that measure aver…