Person-centric Visual Grounding
1개 벤치마크 · 논문 4편 · 이 태스크의 논문 보기 →
Benchmarks
Who’s Waldo
Most implemented
To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo
To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo
TubeDETR: Spatio-Temporal Video Grounding with Transformers
Who's Waldo? Linking People Across Text and Images
Papers
To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who’s Waldo dataset. Given an image and a caption, PCVG requires pairing up a person’s name men…
BenchmarkingPerson-centric Visual GroundingSentenceVisual GroundingTubeDETR: Spatio-Temporal Video Grounding with Transformers
We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal …
DecoderLanguage-Based Temporal LocalizationNatural Language Visual Groundingobject-detection+6To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo
We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG requires pairing up a person's name men…
BenchmarkingPerson-centric Visual GroundingSentenceVisual GroundingWho's Waldo? Linking People Across Text and Images
We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which …
Person-centric Visual Grounding