paper-with-me

홈 › Papers

AI Knows What's Wrong But Cannot Fix It: Helicoid Dynamics in Frontier LLMs Under High-Stakes Decisions

2026-03-12 · Alejandro R Jadad arxiv

Large language models perform reliably when their outputs can be checked: solving equations, writing code, retrieving facts. They perform differently when checking is impossible, as when a clinician chooses an irreversible treatment on incomplete data, or an investor commits capital under fundamental uncertainty. Helicoid dynamics is the name given to a specific failure regime in that second domain: a system engages competently, drifts into error, accurately names what went wrong, then reproduces the same pattern at a higher level of sophistication, recognizing it is looping and continuing nonetheless. This prospective case series documents that regime across seven leading systems (Claude, ChatGPT, Gemini, Grok, DeepSeek, Perplexity, Llama families), tested across clinical diagnosis, investment evaluation, and high-consequence interview scenarios. Despite explicit protocols designed to sustain rigorous partnership, all exhibited the pattern. When confronted with it, they attributed its persistence to structural factors in their training, beyond what conversation can reach. Under high stakes, when being rigorous and being comfortable diverge, these systems tend toward comfort, becoming less reliable precisely when reliability matters most. Twelve testable hypotheses are proposed, with implications for agentic AI oversight and human-AI collaboration. The helicoid is tractable. Identifying it, naming it, and understanding its boundary conditions are the necessary first steps toward LLMs that remain trustworthy partners precisely when the decisions are hardest and the stakes are highest.

📄 PDF Abstract BibTeX arXiv:2603.11559

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

2026-08-17 · Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding 외 arxiv

Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, w…

Printed helicoids with embedded air channels make sensorized segments for soft continuum robots

2026-02-26 · Annan Zhang, Hanna Matusik, Miguel Flores-Acton, Emily R. Sologuren 외 arxiv

Soft robots enable safe, adaptive interaction with complex environments but remain difficult to sense and control due to their highly deformable structures. Architected soft materials such as helicoid lattices offer tuna…

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

2026-05-25 · Ziyi Ding, Xiao-Ping Zhang arxiv

We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed linear probe recovers the evidence-supported answer from the model's h…

Periodic recurrent waves of Covid-19 epidemics and vaccination campaign

2022-01-18 · Gaetano Campi, Antonio Bianconi

While understanding of periodic recurrent waves of Covid-19 epidemics would aid to combat the pandemics, quantitative analysis of data over a two years period from the outbreak, is lacking. The complexity of Covid-19 rec…

Rift: A Conflict Signature for Deception in Language Models

2026-06-15 · Petr Nyoma arxiv

A model that lies while knowing the truth is the central case ELK cannot handle with behavioral evaluation alone. We ask whether such deception leaves an internal signature distinguishing it from honest error. Our key mo…