paper-with-me

Papers Referring Expression Segmentation

“Referring Expression Segmentation” 태그가 달린 논문 164편 · 필터 해제

DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation

2026-08-24 · Yujie Qi, Luyan Zhang arxiv

Referring Expression Segmentation (RES) aims to generate a pixel-level mask for the object specified by a language expression. Recent methods based on multimodal large language models (MLLMs) often rely on one-pass coord…

Referring Expression SegmentationVisual Localization

Falcon Perception-HD: High Density Perception via Reinforcement Learning

2026-08-19 · Sofian Chaybouti, Yasser Dahou, Ngoc Dung Huynh, Reda Alami 외 arxiv

Autoregressive perception models trained to localize visual entities under the open-vocabulary setting are mostly trained using Supervised fine-tuning (SFT) with maximum likelihood, yet it optimizes a proxy objective (pe…

Referring Expression SegmentationReinforcement Learning

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

2026-07-20 · Dingyun Zhang, Lixue Gong, Wei Liu hf

In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting vide…

Referring Expression SegmentationImage Editing

FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation

2026-05-28 · Zekang Zhang, Guangyu Gao, Youyun Tang, ChengJing Wu 외 arxiv

LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we identify a persistent failure mode in current propose-then-select pipeli…

Referring Expression Segmentation

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

2026-05-27 · Runlong Cao, Ying Zang, Chuanwei Zhou, Tianrun Chen 외 arxiv

Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploitin…

Referring Expression Segmentation

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

2026-05-08 · Yuan Yao, Qiushi Yang, Humen Zhong, Jiangning Wei 외 arxiv

Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit strong open-world visual grounding, but th…

Referring Expression SegmentationVisual Grounding

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

2026-04-09 · Weiming Zhang, Dingwen Xiao, Songyue Guo, Guangyu Xiang 외 arxiv

Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision and language understanding. Existing RES methods, however, rely heavily …

Referring Expression Segmentation

Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation

2026-03-19 · Yuchen Li, Amanmeet Garg, Shalini Chaudhuri, Rui Zhao 외 arxiv

Large Vision Language Models (LVLMs) excel at semantic understanding but struggle with fine grained spatial grounding, as the model must implicitly infer complex geometry without ever producing a spatial interpretation. …

Referring Expression SegmentationSemantic SegmentationSpatial Reasoning

SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation

2026-03-18 · Wei Tang, Xuejing Liu, Yanpeng Sun, Zechao Li arxiv

The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Expression Segmentation (RES). Toward this…

Referring Expression SegmentationImage Segmentation

Hierarchical Collaborative Fusion for 3D Instance-aware Referring Expression Segmentation

2026-03-06 · Keshen Zhou, Runnan Chen, Mingming Gong, Tongliang Liu arxiv

Generalised 3D Referring Expression Segmentation (3D-GRES) localizes objects in 3D scenes based on natural language, even when descriptions match multiple or zero targets. Existing methods rely solely on sparse point clo…

Referring Expression SegmentationPoint Clouds

3D-DRES: Detailed 3D Referring Expression Segmentation

2026-03-03 · Qi Chen, Changli Wu, Jiayi Ji, Yiwei Ma 외 arxiv

Current 3D visual grounding tasks only process sentence level detection or segmentation, which critically fails to leverage the rich compositional contextual reasonings within natural language expressions. To address thi…

Referring Expression SegmentationVisual Grounding

ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation

2026-01-23 · Yihao Wang, Jusheng Zhang, Ziyi Tang, Keze Wang 외 arxiv

Referring Expression Segmentation (RES) is a core vision-language segmentation task that enables pixel-level understanding of targets via free-form linguistic expressions, supporting critical applications such as human-r…

Referring Expression SegmentationVisual Reasoning

OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding

2026-01-14 · Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng Sun arxiv

We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi-view im…

Referring Expression SegmentationScene Understanding

MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation

2026-01-11 · Changli Wu, Haodong Wang, Jiayi Ji, Yutian Yao 외 arxiv

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with only a few sparse RGB views and strict la…

Referring Expression SegmentationPoint Clouds

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

2025-09-22 · Ye Liu, Zongyang Ma, Junfu Pu, Zhongang Qi 외 arxiv

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conver…

Referring Expression SegmentationQuestion AnsweringVisual Reasoning

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

2025-08-30 · Boyi Li, Ce Zhang, Richard M. Timmerman, Wenxuan Bao arxiv

The emergence of vision language models (VLMs) bridges the gap between vision and language, enabling multimodal understanding beyond traditional visual-only deep learning models. However, transferring VLMs from the natur…

Referring Expression SegmentationSemantic SegmentationPrompt EngineeringImage Segmentation

VoCap: Video Object Captioning and Segmentation from Any Prompt

2025-08-29 · Jasper Uijlings, Xingyi Zhou, Xiuye Gu, Arsha Nagrani 외 arxiv

Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose VoCap, a flexible video model that cons…

Semi-Supervised Video Object SegmentationReferring Expression Segmentation

Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

2025-08-06 · Jingchao Wang, Zhijian Wu, Dingjiang Huang, Yefeng Zheng 외 arxiv

Reference Expression Segmentation (RES) aims to segment image regions specified by referring expressions and has become popular with the rise of multimodal large models (MLLMs). While MLLMs excel in semantic understandin…

Referring Expression Segmentation

Advancing Visual Large Language Model for Multi-granular Versatile Perception

2025-07-22 · Wentao Xiang, Haoxian Tan, Cong Wei, Yujie Zhong 외 arxiv

Perception is a fundamental task in the field of computer vision, encompassing a diverse set of subtasks that can be systematically categorized into four distinct groups based on two dimensions: prediction type and instr…

Referring Expression SegmentationPanoptic Segmentation

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

2025-07-02 · Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Sen yang 외

Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language in…

Data AugmentationGeneralized Referring Expression SegmentationImage SegmentationReading Comprehension+2
1–20 / 164 다음 →