Referring Expression Segmentation
22개 벤치마크 · 논문 164편 · 이 태스크의 논문 보기 →
Benchmarks
RefCoCo val
RefCOCO+ val
RefCOCO+ test B
RefCOCO+ testA
A2D Sentences
RefCOCOg-val
J-HMDB
DAVIS 2017 (val)
RefCOCOg-test
RefCOCO testA
RefCOCO testB
PhraseCut
RefCOCO
ReferIt
Refer-YouTube-VOS
A2Dre test
CLEVR-Ref+
G-Ref test A
G-Ref test B
G-Ref val
Most implemented
Image Segmentation Using Text and Image Prompts
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding
Segmentation from Natural Language Expressions
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning
Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Papers
DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation
Referring Expression Segmentation (RES) aims to generate a pixel-level mask for the object specified by a language expression. Recent methods based on multimodal large language models (MLLMs) often rely on one-pass coord…
Referring Expression SegmentationVisual LocalizationFalcon Perception-HD: High Density Perception via Reinforcement Learning
Autoregressive perception models trained to localize visual entities under the open-vocabulary setting are mostly trained using Supervised fine-tuning (SFT) with maximum likelihood, yet it optimizes a proxy objective (pe…
Referring Expression SegmentationReinforcement LearningFlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting vide…
Referring Expression SegmentationImage EditingFlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we identify a persistent failure mode in current propose-then-select pipeli…
Referring Expression SegmentationLearning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation
Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploitin…
Referring Expression SegmentationQwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit strong open-world visual grounding, but th…
Referring Expression SegmentationVisual Grounding