paper-with-me

홈 › Papers

LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight

2025-11-25 · Yunze Man, Shihao Wang, Guowen Zhang, Johan Bjorck, Zhiqi Li, Liang-Yan Gui, Jim Fan, Jan Kautz, Yu-Xiong Wang, Zhiding Yu arxiv

To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM toolbox. We present LocateAnything3D, a VLM-native recipe that casts 3D detection as a next-token prediction problem. The key is a short, explicit Chain-of-Sight (CoS) sequence that mirrors how human reason from images: find an object in 2D, then infer its distance, size, and pose. The decoder first emits 2D detections as a visual chain-of-thought, then predicts 3D boxes under an easy-to-hard curriculum: across objects, a near-to-far order reduces early ambiguity and matches ego-centric utility; within each object, a center-from-camera, dimensions, and rotation factorization ranks information by stability and learnability. This VLM-native interface preserves open-vocabulary and visual-prompting capability without specialized heads. On the challenging Omni3D benchmark, our model achieves state-of-the-art results, with 38.90 AP_3D, surpassing the previous best by +13.98 absolute improvement even when the baseline is given ground-truth 2D boxes. It also generalizes zero-shot to held-out categories with strong robustness. By turning 3D detection into a disciplined next-token problem, LocateAnything3D offers a practical foundation for models to perceive in 3D.

📄 PDF Abstract BibTeX arXiv:2511.20648

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

2026-05-26 · Shihao Wang, Shilong Liu, Yuanguo Kuang, Xinyu Wei 외 arxiv

Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are learned and decoded largely independently…

Visual Grounding

Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight

2024-07-22 · Ziyuan Huang, Kaixiang Ji, Biao Gong, Zhiwu Qing 외

This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visua…

GenAI-Driven Approach to RISC-V Supply Chain Exploration

2026-05-13 · Nenad Petrovic, Andre Schamschurko, Yingjie Xu, Alois Knoll arxiv

This paper presents an LLM-empowered workflow for RISC-V supply chain analysis, integrating Vision-Language Models (VLMs) and Model-Driven Engineering (MDE) to enable comprehensive, multimodal data-driven insights. The p…

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

2026-08-04 · Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang 외 arxiv

Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models…

Visual Grounding

Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards

2025-09-23 · Honghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang 외 arxiv

Chain of thought reasoning has demonstrated remarkable success in large language models, yet its adaptation to vision-language reasoning remains an open challenge with unclear best practices. Existing attempts typically …

Reinforcement LearningMultimodal Reasoning