paper-with-me

홈 › Papers

Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning

2020-09-30 · ECCV 2020 8 · Xiangxi Shi, Xu Yang, Jiuxiang Gu, Shafiq Joty, Jianfei Cai

Change Captioning is a task that aims to describe the difference between images with natural language. Most existing methods treat this problem as a difference judgment without the existence of distractors, such as viewpoint changes. However, in practice, viewpoint changes happen often and can overwhelm the semantic difference to be described. In this paper, we propose a novel visual encoder to explicitly distinguish viewpoint changes from semantic changes in the change captioning task. Moreover, we further simulate the attention preference of humans and propose a novel reinforcement learning process to fine-tune the attention directly with language evaluation rewards. Extensive experimental results show that our method outperforms the state-of-the-art approaches by a large margin in both Spot-the-Diff and CLEVR-Change datasets.

📄 PDF Abstract BibTeX arXiv:2009.14352

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Semantic Part Detection via Matching: Learning to Generalize to Novel Viewpoints from Limited Training Data

2018-11-28 · ICCV 2019 10 · Yutong Bai, Qing Liu, Lingxi Xie, Weichao Qiu 외

Detecting semantic parts of an object is a challenging task in computer vision, particularly because it is hard to construct large annotated datasets due to the difficulty of annotating semantic parts. In this paper we p…

ClusteringObjectSemantic Part Detection

GAN-based Pose-aware Regulation for Video-based Person Re-identification

2019-03-27 · Alessandro Borgia, Yang Hua, Elyor Kodirov, Neil M. Robertson

Video-based person re-identification deals with the inherent difficulty of matching unregulated sequences with different length and with incomplete target pose/viewpoint structure. Common approaches operate either by red…

Person Re-IdentificationVideo-Based Person Re-Identification

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

2026-07-22 · Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao 외 arxiv

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference i…

Visual LocalizationPose Estimation

Person Re-identification Based on Color Histogram and Spatial Configuration of Dominant Color Regions

2014-11-13 · Kwangchol Jang, Sokmin Han, In-Song Kim

There is a requirement to determine whether a given person of interest has already been observed over a network of cameras in video surveillance systems. A human appearance obtained in one camera is usually different fro…

Person Re-Identification

Comparative Evaluation of Hand-Crafted and Learned Local Features

2017-07-01 · Conference on Computer Vision and Pattern Recognition 2017 7 · Johannes L. Sch¨onberger, Hans Hardmeier, Torsten Sattler, Marc Pollefeys

Matching local image descriptors is a key step in many computer vision applications. For more than a decade,hand-crafted descriptors such as SIFT have been used for this task. Recently, multiple new descriptors learned f…

Image RetrievalRetrieval