paper-with-me

홈 › Papers

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

2026-06-24 · Jaeyong Ko, Pilsung Kang, Yukyung Lee arxiv

Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail. Prior work analyzes failure at the step, chunk, or sentence level, or at tokens where failure has already occurred. Neither identifies the precise token that triggers the shift toward failure. We introduce the cliff token, a token where the token-wise potential drops significantly under an adaptive threshold that scales with the local token-wise potential, based on a one-sided two-proportion z-test. Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to between 0.71 and 1.00. We further introduce a cliff taxonomy of deterministic, uncertain, and sampled-off cliffs, defined by greedy choice and token entropy. Each type has distinct probabilistic characteristics, and the taxonomy generalizes across model scales. Finally, we validate the taxonomy via single-token preference optimization at cliff positions (Cliff-DPO). Trained on GSM8K, Cliff-DPO improves accuracy across benchmarks by up to +6.6. Optimizing at uncertain and sampled-off cliffs improves reasoning, while deterministic cliffs do not.

📄 PDF Abstract BibTeX arXiv:2606.25524

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

ContractBench: Can LLM Agents Preserve Observation Contracts?

2026-05-17 · Jicheng Wang, Yifeng He, Zili Wang, Hanwen Xing 외 arxiv

Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation contracts: artifacts whose later use is constrained by the external syst…

Wrapping trust for interoperability. A study of wrapped tokens

2021-09-14 · Giulio Caldarelli

As known, blockchains are traditionally blind to the real world. This implies the reliance on third parties called oracles when extrinsic data is needed for smart contracts. However, reintroducing trust and single point …

When Molecular Similarity Works: Property Cliffs Reveal Hidden Errors

2026-05-17 · Di Hu, Kun Li, Haojie Rao, Longtao Hu 외 arxiv

Accurate prediction of molecular properties underpins drug discovery and material design, yet even state-of-the-art models remain vulnerable to localized failure modes that aggregate metrics cannot detect. The places whe…

Drug Discovery

Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?

2025-10-07 · Qingyu Yin, Chak Tou Leong, Linyi Yang, Wenxuan Huang 외 arxiv

Large reasoning models (LRMs) with multi-step reasoning capabilities have shown remarkable problem-solving abilities, yet they exhibit concerning safety vulnerabilities that remain poorly understood. In this work, we inv…

LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

2026-01-29 · Xunkai Li, Zhengyu Wu, Zekai Chen, Henan Sun 외 arxiv

Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data…

Graph Learning