paper-with-me

홈 › Papers

Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction

2025-10-23 · Suchir Salhan, Hongyi Gu, Donya Rooein, Diana Galvan-Sosa, Gabrielle Gaudeau, Andrew Caines, Zheng Yuan, Paula Buttery arxiv

Multi-turn dialogues between a child and a caregiver are characterized by a property called contingency - that is, prompt, direct, and meaningful exchanges between interlocutors. We introduce ContingentChat, a teacher-student framework that benchmarks and improves multi-turn contingency in a BabyLM trained on 100M words. Using a novel alignment dataset for post-training, BabyLM generates responses that are more grammatical and cohesive. Experiments with adaptive teacher decoding strategies show limited additional gains. ContingentChat demonstrates the benefits of targeted post-training for dialogue quality and indicates that contingency remains a challenging goal for BabyLMs.

📄 PDF Abstract BibTeX arXiv:2510.20411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

2026-06-16 · Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang 외 arxiv

Knowledge distillation transfers a teacher's competence to a small student but is brittle in the small-student regime: forcing the student to imitate logits from a much larger teacher concentrates it on the teacher's sha…

Knowledge DistillationReinforcement Learning

Investigating the Zone of Proximal Development of Language Models for In-Context Learning

2025-02-10 · Peng Cui, Mrinmaya Sachan

In this paper, we introduce a learning analytics framework to analyze the in-context learning (ICL) behavior of large language models (LLMs) through the lens of the Zone of Proximal Development (ZPD), an established theo…

In-Context Learning

Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty

2023-08-03 · Inar Timiryasov, Jean-Loup Tastet

We present our submission to the BabyLM challenge, whose goal was to improve the sample efficiency of language models. We trained an ensemble consisting of a GPT-2 and small LLaMA models on the developmentally-plausible,…

Knowledge Distillation

Can training neural language models on a curriculum with developmentally plausible data improve alignment with human reading behavior?

2023-11-30 · Aryaman Chobey, Oliver Smith, Anzi Wang, Grusha Prasad

The use of neural language models to model human behavior has met with mixed success. While some work has found that the surprisal estimates from these models can be used to predict a wide range of human neural and behav…

Sentence

Closed-loop Teaching via Demonstrations to Improve Policy Transparency

2024-04-01 · Michael S. Lee, Reid Simmons, Henny Admoni

Demonstrations are a powerful way of increasing the transparency of AI policies. Though informative demonstrations may be selected a priori through the machine teaching paradigm, student learning may deviate from the pre…