paper-with-me

홈 › Papers

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

2026-05-11 · Hao Liu, Jicheng Liu arxiv

A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus from 1,951 prime-knot prototypes (crossing numbers 3 to 19) with a protocol whose answers are checked against Regina's canonical knot signature. Its 14 tasks span four families, equivalence judgment, move prediction, identification, and cross-modal grounding; an image-versus-symbol split locates failures along the perception-operation gap. We score Claude Opus 4.7 and GPT-5, each with and without thinking, under a 64K output-token budget matched on both vendors. Across 56 (task, model) cases, 15 sit at or below a random baseline and 8 of 14 tasks have a best score under 1.5x random. On diagram-to-symbol transcription, no model produces a strictly correct string, and permissive Regina decoding recovers the knot in 0 to 4 of 100 items. Thinking-mode reasoning lifts overall accuracy by 1.65 points for Claude and 9.25 points for GPT-5, narrowing the gap only modestly. Read together, the four families suggest current vision-language models hold features of a diagram but lack apparatus to simulate moves on those features.

📄 PDF Abstract BibTeX arXiv:2605.09900

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Knot So Simple: A Minimalistic Environment for Spatial Reasoning

2025-05-23 · Zizhao Chen, Yoav Artzi

We propose KnotGym, an interactive environment for complex, spatial reasoning and manipulation. KnotGym includes goal-oriented rope manipulation tasks with varying levels of complexity, all requiring acting from pure ima…

Model Predictive ControlSpatial Reasoning

The unknotting number, hard unknot diagrams, and reinforcement learning

2024-09-13 · Taylor Applebaum, Sam Blackwell, Alex Davies, Thomas Edlich 외

We have developed a reinforcement learning agent that often finds a minimal sequence of unknotting crossing changes for a knot diagram with up to 200 crossings, hence giving an upper bound on the unknotting number. We ha…

reinforcement-learningReinforcement Learning

RL unknotter, hard unknots and unknotting number

2026-03-09 · Anne Dranowski, Yura Kabkov, Daniel Tubbenhauer arxiv

We develop a reinforcement learning pipeline for simplifying knot diagrams. A trained agent learns move proposals and a value heuristic for navigating Reidemeister moves. The pipeline applies to arbitrary knots and links…

Reinforcement Learning

AlphaFold predicts the most complex protein knot and composite protein knots

2022-07-15 · Maarten A. Brems, Robert Runkel, Todd O. Yeates, Peter Virnau

The computer artificial intelligence system AlphaFold has recently predicted previously unknown three-dimensional structures of thousands of proteins. Focusing on the subset with high-confidence scores, we algorithmicall…

Learning to Unknot

2020-10-28 · Sergei Gukov, James Halverson, Fabian Ruehle, Piotr Sułkowski

We introduce natural language processing into the study of knot theory, as made natural by the braid word representation of knots. We study the UNKNOT problem of determining whether or not a given knot is the unknot. Aft…

Binary ClassificationReinforcement Learning (RL)