paper-with-me

Papers

BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models

2026-01-26 · Kaustubh D. Dhole arxiv

Traditional evaluations of reasoning capabilities of language models are dominated by adult-centric benchmarks that presuppose broad world knowledge, complex instruction following, and mature pragmatic competence. These assumptions are mismatched to baby language models trained on developmentally plausible input such as child-directed speech and early-childhood narratives, and they obscure which reasoning abilities (if any) emerge under such constraints. We introduce BabyReasoningBench, a GPT-5.2 generated benchmark of 19 reasoning tasks grounded in classic paradigms from developmental psychology, spanning theory of mind, analogical and relational reasoning, causal inference and intervention selection, and core reasoning primitives that are known to be confounded by memory and pragmatics. We find that two GPT-2 based baby language models (pretrained on 10M and 100M of child-directed speech text) show overall low but uneven performance, with dissociations across task families: scaling improves several causal and physical reasoning tasks, while belief attribution and pragmatics-sensitive tasks remain challenging. BabyReasoningBench provides a developmentally grounded lens for analyzing what kinds of reasoning are supported by child-like training distributions, and for testing mechanistic hypotheses about how such abilities emerge.

📄 PDF Abstract BibTeX arXiv:2601.18933

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingRelational ReasoningCausal Inference

Similar Papers 제목 키워드 기반

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

2025-12-11 · Shengao Wang, Wenqi Wang, Zecheng Wang, Max Whitton 외 arxiv

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-lan…

Spatial Reasoning

Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

2023-01-27 · Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams 외

We present the call for papers for the BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus. This shared task is intended for participants with an interest in small scale language modeling…

Language AcquisitionLanguage ModelingLanguage ModellingNatural Language Understanding

Do Syntactic Categories Help in Developmentally Motivated Curriculum Learning for Language Models?

2025-11-11 · Arzu Burcu Güven, Anna Rogers, Rob van der Goot arxiv

We examine the syntactic properties of BabyLM corpus, and age-groups within CHILDES. While we find that CHILDES does not exhibit strong syntactic differentiation by age, we show that the syntactic knowledge about the tra…

Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity

2026-04-22 · Pranava Madhyastha, Dagmar Adamcova arxiv

We investigate the integration of human-like working memory constraints into the Transformer architecture and implement several cognitively inspired attention variants, including fixed-width windows based and temporal de…

A Small Math Model: Recasting Strategy Choice Theory in an LLM-Inspired Architecture

2025-09-28 · Roussel Rahman, Jeff Shrager arxiv

Strategy Choice Theory (SCT; Siegler and Shrager, 1984; Siegler, 2000) explains important aspects of children's arithmetic learning based upon principles including learning from developmentally naturalistic data, probabi…

Mathematical Reasoning