paper-with-me

Papers

StAR: Segment Anything Reasoner

2026-03-15 · Seokju Yun, Dongheon Lee, Noori Bae, Jaesung Jun, Chanseul Cho, Youngmin Ro arxiv

As AI systems are being integrated more rapidly into diverse and complex real-world environments, the ability to perform holistic reasoning over an implicit query and an image to localize a target is becoming increasingly important. However, recent reasoning segmentation methods fail to sufficiently elicit the visual reasoning capabilities of the base mode. In this work, we present Segment Anything Reasoner (StAR), a comprehensive framework that refines the design space from multiple perspectives-including parameter-tuning scheme, reward functions, learning strategies and answer format-and achieves substantial improvements over recent baselines. In addition, for the first time, we successfully introduce parallel test-time scaling to the segmentation task, pushing the performance boundary even further. To extend the scope and depth of reasoning covered by existing benchmark, we also construct the ReasonSeg-X, which compactly defines reasoning types and includes samples that require deeper reasoning. Leveraging this dataset, we train StAR with a rollout-expanded selective-tuning approach to activate the base model's latent reasoning capabilities, and establish a rigorous benchmark for systematic, fine-grained evaluation of advanced methods. With only 5k training samples, StAR achieves significant gains over its base counterparts across extensive benchmarks, demonstrating that our method effectively brings dormant reasoning competence to the surface.

📄 PDF Abstract BibTeX arXiv:2603.14382

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Convex Combination Star Shape Prior for Data-driven Image Semantic Segmentation

2025-01-01 · CVPR 2025 1 · Xinyu Zhao, Jun Xie, Shengzhe Chen, Jun Liu

Multi-center star shape is a prevalent object shape feature, which has proven effective in model-based image segmentation methods. However, the shape field function induced by the multi-center star shape is non-smoot…

Image SegmentationSegmentationSemantic Segmentation

Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation

2025-07-04 · Tao Tang, Shijie Xu, Yiting Wu, Zhixiang Lu

The clinical utility of deep learning models for medical image segmentation is severely constrained by their inability to generalize to unseen domains. This failure is often rooted in the models learning spurious correla…

AnatomyDisentanglementImage SegmentationMedical Image Segmentation+1

SAM.MD: Zero-shot medical image segmentation capabilities of the Segment Anything Model

2023-04-10 · Saikat Roy, Tassilo Wald, Gregor Koehler, Maximilian R. Rokuss 외

Foundation models have taken over natural language processing and image generation domains due to the flexibility of prompting. With the recent introduction of the Segment Anything Model (SAM), this prompt-driven paradig…

Image GenerationImage SegmentationMedical Image SegmentationOrgan Segmentation+3

DeiSAM: Segment Anything with Deictic Prompting

2024-02-21 · Hikaru Shindo, Manuel Brack, Gopika Sudhakaran, Devendra Singh Dhami 외

Large-scale, pre-trained neural networks have demonstrated strong capabilities in various tasks, including zero-shot image segmentation. To identify concrete objects in complex scenes, humans instinctively rely on deicti…

Image SegmentationSegmentationSemantic Segmentation

SAMAug: Point Prompt Augmentation for Segment Anything Model

2023-07-03 · Haixing Dai, Chong Ma, Zhiling Yan, Zhengliang Liu 외

This paper introduces SAMAug, a novel visual point augmentation method for the Segment Anything Model (SAM) that enhances interactive image segmentation performance. SAMAug generates augmented point prompts to provide mo…

Image SegmentationmodelPrompt EngineeringSegmentation+1