paper-with-me

홈 › Papers

Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection

2026-01-10 · Minghui Jia, Qichao Zhang, Ali Luo, Linjing Li, Shuo Ye, Hailing Lu, Wen Hou, Dongbin Zhao arxiv

Due to the limited generalization and interpretability of deep learning classifiers, The final vetting of rare celestial object candidates still relies on expert visual inspection--a manually intensive process. In this process, astronomers leverage specialized tools to analyze spectra and construct reliable catalogs. However, this practice has become the primary bottleneck, as it is fundamentally incapable of scaling with the data deluge from modern spectroscopic surveys. To bridge this gap, we propose Spec-o3, a tool-augmented vision-language agent that performs astronomer-aligned spectral inspection via interleaved multimodal chain-of-thought reasoning. Spec-o3 is trained with a two-stage post-training recipe: cold-start supervised fine-tuning on expert inspection trajectories followed by outcome-based reinforcement learning on rare-type verification tasks. Evaluated on five rare-object identification tasks from LAMOST, Spec-o3 establishes a new State-of-the-Art, boosting the macro-F1 score from 28.3 to 76.5 with a 7B parameter base model and outperforming both proprietary VLMs and specialized deep models. Crucially, the agent demonstrates strong generalization to unseen inspection tasks across survey shifts (from LAMOST to SDSS/DESI). Expert evaluations confirm that its reasoning traces are coherent and physically consistent, supporting transparent and trustworthy decision-making. Code, data, and models are available at https://github.com/Maxwell-Jia/spec-o3.

📄 PDF Abstract BibTeX arXiv:2601.06498

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

2026-06-02 · Anjie Liu, Yan Song, Zhixun Chen, Ziqin Gong 외 arxiv

Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing every proposed tool call is costly and sometimes unnecessary. We study t…

Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution

2025-11-18 · N Dinesh Reddy, Dylan Snyder, Lona Kiragu, Mirajul Mohin 외 arxiv

We introduce Orion, a visual agent that integrates vision-based reasoning with tool-augmented execution to achieve powerful, precise, multi-step visual intelligence across images, video, and documents. Unlike traditional…

Panoptic SegmentationObject DetectionVisual Reasoning

TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

2026-08-26 · Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi 외 arxiv

Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agentic retrieval-augmen…

Answer GenerationVideo Captioning

EAA: Automating materials characterization with vision language model agents

2026-02-17 · Ming Du, Yanqi Luo, Srutarshi Banerjee, Michael Wojcik 외 arxiv

We present Experiment Automation Agents (EAA), a vision-language-model-driven agentic system designed to automate complex experimental microscopy workflows. EAA integrates multimodal reasoning, tool-augmented action, and…

Multimodal Reasoning

Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning

2026-06-11 · Changye Li, Meng Lu, Yi Wu, Ligeng Zhu arxiv

While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limi…

Spatial Reasoning