paper-with-me

Person-centric Visual Grounding

1개 벤치마크 · 논문 4편 · 이 태스크의 논문 보기 →

Benchmarks

Who’s Waldo

결과 1개

Most implemented

Papers

To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo

2022-05-01 · ACL 2022 5 · Yiran Luo, Pratyay Banerjee, Tejas Gokhale, Yezhou Yang 외

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who’s Waldo dataset. Given an image and a caption, PCVG requires pairing up a person’s name men…

BenchmarkingPerson-centric Visual GroundingSentenceVisual Grounding

TubeDETR: Spatio-Temporal Video Grounding with Transformers

2022-03-30 · CVPR 2022 1 · Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 외

We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal …

DecoderLanguage-Based Temporal LocalizationNatural Language Visual Groundingobject-detection+6

To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo

2022-03-30 · Yiran Luo, Pratyay Banerjee, Tejas Gokhale, Yezhou Yang 외

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG requires pairing up a person's name men…

BenchmarkingPerson-centric Visual GroundingSentenceVisual Grounding

Who's Waldo? Linking People Across Text and Images

2021-08-16 · ICCV 2021 10 · Claire Yuqing Cui, Apoorv Khandelwal, Yoav Artzi, Noah Snavely 외

We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which …

Person-centric Visual Grounding