paper-with-me

홈 › Papers

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

2026-06-15 · Shiyang Chen arxiv

LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a lens concurrent work sets aside -- the model's attention to labeled tool-definition segments. On real BFCL failures, by per-candidate attention argmax the model attends most to the correct tool 80% of the time (vs. 21% chance), and the gold is the under-attended segment on only 10%: it looks at the right tool and still picks wrong. This directly refutes the intuitive "crowded-harness / lost-in-the-middle" explanation: the failure is at the decision readout, not the harness, and we pin it there three ways. (1) Input vs. readout: repairing the prompt (reordering or duplicating the gold tool) recovers <=23% of failures, while readout-side interventions recover 59-91%. (2) Representation-invariance: two gold-pointed interventions in different representations -- an additive attention-logit bias and a residual-stream steering vector -- recover largely the same failures (per-task Jaccard 0.865 pooled, 0.79-0.91 per model), so the bottleneck is localized to the readout independent of which representation is poked. (3) A training-free, gold-free selector: per-segment attention closes most of the gold-free-vs-oracle gap on BFCL (+11.9 pts pooled function-name selection vs. +17.9-pt oracle headroom) and adds +14.9 pts on Seal-Tools; every model positive (exact McNemar p<=8e-4 each). Scopes differ: the causal attention-bias dose-response is bidirectional and monotonic on 10 mask-honoring models (3-32B), the full 0.5-32B span carrying only the correlational diagnostic; the deployable selector is evaluated on 5 single-turn models and does not yet transfer to a multi-turn loop.

📄 PDF Abstract BibTeX arXiv:2606.16364

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Instance Segmentation for Autonomous Log Grasping in Forestry Operations

2022-03-03 · Jean-Michel Fortin, Olivier Gamache, Vincent Grondin, François Pomerleau 외

Wood logs picking is a challenging task to automate. Indeed, logs usually come in cluttered configurations, randomly orientated and overlapping. Recent work on log picking automation usually assume that the logs' pose is…

Inductive BiasInstance SegmentationSemantic Segmentation

First arrival picking using U-net with Lovasz loss and nearest point picking method

2021-04-06 · Pengyu Yuan, Wenyi Hu, Xuqing Wu, Jiefu Chen 외

We proposed a robust segmentation and picking workflow to solve the first arrival picking problem for seismic signal processing. Unlike traditional classification algorithm, image segmentation method can utilize the loca…

Contour DetectionImage SegmentationSegmentationSemantic Segmentation

Automatic Period Segmentation of Oral French

2020-05-01 · LREC 2020 5 · Natalia Kalashnikova, Lo{\"\i}c Grobol, Iris Eshkol-Taravella, Fran{\c{c}}ois Delafontaine

Natural Language Processing in oral speech segmentation is still looking for a minimal unit to analyze. In this work, we present a comparison of two automatic segmentation methods of macro-syntactic periods which allows …

Segmentation

ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics

2025-12-10 · Donato Caramia, Florian T. Pokorny, Giuseppe Triggiani, Denis Ruffino 외 arxiv

Occlusions in robotic bin picking compromise accurate and reliable grasp planning. We present ViTA-Seg, a class-agnostic Vision Transformer framework for real-time amodal segmentation that leverages global attention to r…

Computational Efficiency

Multi-Task Deep Networks for Depth-Based 6D Object Pose and Joint Registration in Crowd Scenarios

2018-06-11 · Juil Sock, Kwang In Kim, Caner Sahin, Tae-Kyun Kim

In bin-picking scenarios, multiple instances of an object of interest are stacked in a pile randomly, and hence, the instances are inherently subjected to the challenges: severe occlusion, clutter, and similar-looking di…

3D Pose EstimationMulti-Task LearningObjectPose Estimation