Papers Person-centric Visual Grounding
“Person-centric Visual Grounding” 태그가 달린 논문 4편 · 필터 해제
To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who’s Waldo dataset. Given an image and a caption, PCVG requires pairing up a person’s name men…
BenchmarkingPerson-centric Visual GroundingSentenceVisual GroundingTubeDETR: Spatio-Temporal Video Grounding with Transformers
We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal …
DecoderLanguage-Based Temporal LocalizationNatural Language Visual Groundingobject-detection+6To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG requires pairing up a person's name men…
BenchmarkingPerson-centric Visual GroundingSentenceVisual GroundingWho's Waldo? Linking People Across Text and Images
We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which …
Person-centric Visual Grounding