paper-with-me

홈 › Papers

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

2026-08-13 · Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel arxiv

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.

📄 PDF Abstract BibTeX arXiv:2608.13545

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation

2026-02-12 · Bowei He, Yankai Chen, Xiaokun Zhang, Linghe Kong 외 arxiv

Knowledge distillation from Large Language Models (LLMs) to smaller models has emerged as a critical technique for deploying efficient AI systems. However, current methods for distillation via synthetic data lack pedagog…

Knowledge Distillation

PEMUTA: Pedagogically-Enriched Multi-Granular Undergraduate Thesis Assessment

2025-07-25 · Jialu Zhang, Qingyang Sun, Qianyi Wang, Weiyi Zhang 외 arxiv

The undergraduate thesis (UGTE) plays an indispensable role in assessing a student's cumulative academic development throughout their college years. Although large language models (LLMs) have advanced education intellige…

How to Build an AI Tutor That Can Adapt to Any Course Using Knowledge Graph-Enhanced Retrieval-Augmented Generation (KG-RAG)

2023-11-29 · Chenxi Dong, Yimin Yuan, Kan Chen, Shupei Cheng 외

This paper introduces KG-RAG (Knowledge Graph-enhanced Retrieval-Augmented Generation), a novel framework that addresses two critical challenges in LLM-based tutoring systems: information hallucination and limited course…

HallucinationKnowledge GraphsLanguage ModelingLanguage Modelling+5

You Get what You Annotate: A Pedagogically Annotated Corpus of Coursebooks for Swedish as a Second Language

2014-11-01 · WS 2014 11 · Elena Volodina, Ildik{\'o} Pil{\'a}n, Stian R{\o}dven Eide, Hannes Heidarsson
Language Acquisition

Impact of Guidance and Interaction Strategies for LLM Use on Learner Performance and Perception

2023-10-13 · Harsh Kumar, Ilya Musabirov, Mohi Reza, Jiakai Shi 외

Personalized chatbot-based teaching assistants can be crucial in addressing increasing classroom sizes, especially where direct teacher presence is limited. Large language models (LLMs) offer a promising avenue, with inc…

Chatbot