paper-with-me

홈 › Papers

End-to-end Planner Training for Language Modeling

2024-10-16 · Nathan Cornille, Florian Mai, Jingyuan Sun, Marie-Francine Moens

Through end-to-end training to predict the next token, LLMs have become valuable tools for various tasks. Enhancing their core training in language modeling can improve numerous downstream applications. A successful approach to enhance language modeling uses a separate planning module to predict abstract labels of future sentences and conditions the LM on these predictions. However, this method is non-differentiable, preventing joint end-to-end tuning of the planner with the LM. We propose an effective method to improve this approach by enabling joint fine-tuning of the planner and the LM. We show that a naive way of approximating the gradient of selecting a label via the straight-through estimator is not effective. Instead, we propose to use the predicted label probabilities as mixing weights to condition the LM on a weighted average of label embeddings in a differentiable manner. This not only enables joint fine-tuning of the planner and the LM, but also allows the LM to draw on the full label distribution predicted by the planner, retaining more information. Our experimental results show consistent improvements in perplexity.

📄 PDF Abstract BibTeX arXiv:2410.12492

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Learning to Plan for Language Modeling from Unlabeled Data

2024-03-31 · Nathan Cornille, Marie-Francine Moens, Florian Mai

By training to predict the next token in an unlabeled corpus, large language models learn to perform many tasks without any labeled data. However, their next-token-prediction objective arguably limits their performance i…

Language ModelingLanguage ModellingSelf-Supervised Learning

VDLM: Variable Diffusion LMs via Robust Latent-to-Text Rendering

2026-01-27 · Shuhui Qu arxiv

Autoregressive language models decode left-to-right with irreversible commitments, limiting revision during multi-step reasoning. We propose \textbf{VDLM}, a modular variable diffusion language model that separates seman…

Task-agnostic Pre-training and Task-guided Fine-tuning for Versatile Diffusion Planner

2024-09-30 · Chenyou Fan, Chenjia Bai, Zhao Shan, Haoran He 외

Diffusion models have demonstrated their capabilities in modeling trajectories of multi-tasks. However, existing multi-task planners or policies typically rely on task-specific demonstrations via multi-task imitation, or…

Reinforcement Learning (RL)

Real-World Planning with PDDL+ and Beyond

2024-02-19 · Wiktor Piotrowski, Alexandre Perez

Real-world applications of AI Planning often require a highly expressive modeling language to accurately capture important intricacies of target systems. Hybrid systems are ubiquitous in the real-world, and PDDL+ is the …

DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving

2026-02-06 · Feiyang jia, Lin Liu, Ziying Song, Caiyan Jia 외 arxiv

End-to-end (E2E) autonomous driving has recently attracted increasing interest in unifying Vision-Language-Action (VLA) with World Models to enhance decision-making and forward-looking imagination. However, existing meth…

Autonomous Driving