paper-with-me

홈 › Papers

Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics

2024-05-07 · Hanlin Zhu, Baihe Huang, Shaolun Zhang, Michael Jordan, Jiantao Jiao, Yuandong Tian, Stuart Russell

Auto-regressive large language models (LLMs) show impressive capacities to solve many complex reasoning tasks while struggling with some simple logical reasoning tasks such as inverse search: when trained on '$A \to B$' (e.g., 'Tom is the parent of John'), LLM fails to directly conclude '$B \gets A$' (e.g., 'John is the child of Tom') during inference even if the two sentences are semantically identical, which is known as the 'reversal curse'. In this paper, we theoretically analyze the reversal curse via the training dynamics of (stochastic) gradient descent for two auto-regressive models: (1) a bilinear model that can be viewed as a simplification of a one-layer transformer; (2) one-layer transformers under certain assumptions. Our analysis reveals that for both models, the reversal curse is a consequence of the (effective) model weights 'asymmetry', i.e., the increase of weights from a token $A$ to token $B$ during training does not necessarily cause the increase of the weights from $B$ to $A$, which is caused by the training dynamics under certain choice of loss function and the optimization space of model parameters. Moreover, our analysis can be naturally applied to other logical reasoning tasks such as chain-of-thought (COT), which provides a new perspective different from previous work that focuses on expressivity. Finally, we conduct experiments to validate our theory on multi-layer transformers under different settings. Our code is available at https://github.com/marlo-z/reversal_curse_analysis/.

📄 PDF Abstract BibTeX arXiv:2405.04669

Code (1)

marlo-z/reversal_curse_analysis 공식 구현 pytorch

Tasks

Logical Reasoning

Similar Papers 제목 키워드 기반

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge

2026-02-02 · Xutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi arxiv

Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" -- when trained on forward knowled…

Logical Reasoning

An Analysis and Mitigation of the Reversal Curse

2023-11-13 · Ang Lv, Kaiyi Zhang, Shufang Xie, Quan Tu 외

Recent research observed a noteworthy phenomenon in large language models (LLMs), referred to as the ``reversal curse.'' The reversal curse is that when dealing with two entities, denoted as $a$ and $b$, connected by the…

DenoisingLanguage Modelling

Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure

2025-04-02 · Boshi Wang, Huan Sun

Despite their impressive capabilities, LLMs exhibit a basic generalization failure known as the Reversal Curse, where they struggle to learn reversible factual associations. Understanding why this occurs could help ident…

Arithmetic ReasoningData Augmentation

The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More

2024-06-07 · Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams 외

Today's best language models still struggle with hallucinations: factually incorrect generations, which impede their ability to reliably retrieve information seen during training. The reversal curse, where models cannot …

Information RetrievalRetrieval

Reverse Training to Nurse the Reversal Curse

2024-03-20 · Olga Golovneva, Zeyuan Allen-Zhu, Jason Weston, Sainbayar Sukhbaatar

Large language models (LLMs) have a surprising failure: when trained on "A has a feature B", they do not generalize to "B is a feature of A", which is termed the Reversal Curse. Even when training with trillions of token…