paper-with-me

홈 › Papers

The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner

2025-07-17 · Zhouqi Hua, Wenwei Zhang, Chengqi Lyu, Yuzhe Gu, Songyang Gao, Kuikun Liu, Kai Chen

Length generalization, the ability to solve problems of longer sequences than those observed during training, poses a core challenge of Transformer-based large language models (LLM). Although existing studies have predominantly focused on data-driven approaches for arithmetic operations and symbolic manipulation tasks, these approaches tend to be task-specific with limited overall performance. To pursue a more general solution, this paper focuses on a broader case of reasoning problems that are computable, i.e., problems that algorithms can solve, thus can be solved by the Turing Machine. From this perspective, this paper proposes Turing MAchine Imitation Learning (TAIL) to improve the length generalization ability of LLMs. TAIL synthesizes chain-of-thoughts (CoT) data that imitate the execution process of a Turing Machine by computer programs, which linearly expands the reasoning steps into atomic states to alleviate shortcut learning and explicit memory fetch mechanism to reduce the difficulties of dynamic and long-range data access in elementary operations. To validate the reliability and universality of TAIL, we construct a challenging synthetic dataset covering 8 classes of algorithms and 18 tasks. Without bells and whistles, TAIL significantly improves the length generalization ability as well as the performance of Qwen2.5-7B on various tasks using only synthetic data, surpassing previous methods and DeepSeek-R1. The experimental results reveal that the key concepts in the Turing Machine, instead of the thinking styles, are indispensable for TAIL for length generalization, through which the model exhibits read-and-write behaviors consistent with the properties of the Turing Machine in their attention layers. This work provides a promising direction for future research in the learning of LLM reasoning from synthetic data.

📄 PDF Abstract BibTeX arXiv:2507.13332

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation Learning

Similar Papers 제목 키워드 기반

Imitating Opponent to Win: Adversarial Policy Imitation Learning in Two-player Competitive Games

2022-10-30 · The Viet Bui, Tien Mai, Thanh H. Nguyen

Recent research on vulnerabilities of deep reinforcement learning (RL) has shown that adversarial policies adopted by an adversary agent can influence a target RL agent (victim agent) to perform poorly in a multi-agent e…

Deep Reinforcement LearningImitation LearningMuJoCoReinforcement Learning (RL)

Sequential Causal Imitation Learning with Unobserved Confounders

2022-08-12 · NeurIPS 2021 12 · Daniel Kumor, Junzhe Zhang, Elias Bareinboim

"Monkey see monkey do" is an age-old adage, referring to na\"ive imitation without a deep understanding of a system's underlying mechanics. Indeed, if a demonstrator has access to information unavailable to the imitator …

Decision MakingImitation Learning

State Alignment-based Imitation Learning

2019-11-21 · ICLR 2020 1 · Fangchen Liu, Zhan Ling, Tongzhou Mu, Hao Su

Consider an imitation learning problem that the imitator and the expert have different dynamics models. Most of the current imitation learning methods fail because they focus on imitating actions. We propose a novel stat…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

The Meta-Turing Test

2022-05-11 · Toby Walsh

We propose an alternative to the Turing test that removes the inherent asymmetry between humans and machines in Turing's original imitation game. In this new test, both humans and machines judge each other. We argue that…

Imitator Learning: Achieve Out-of-the-Box Imitation Ability in Variable Environments

2023-10-09 · Xiong-Hui Chen, Junyin Ye, Hang Zhao, Yi-Chen Li 외

Imitation learning (IL) enables agents to mimic expert behaviors. Most previous IL techniques focus on precisely imitating one policy through mass demonstrations. However, in many applications, what humans require is the…

Imitation Learning