paper-with-me

Papers

Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval

2018-08-23 · Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis, Amit K. Roy-Chowdhury

Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great promise by learning deep representations aligned across modalities, most of these methods are plagued by the issue of training with small-scale datasets covering a limited number of images with ground-truth sentences. Moreover, it is extremely expensive to create a larger dataset by annotating millions of images with sentences and may lead to a biased model. Inspired by the recent success of webly supervised learning in deep neural networks, we capitalize on readily-available web images with noisy annotations to learn robust image-text joint representation. Specifically, our main idea is to leverage web images and corresponding tags, along with fully annotated datasets, in training for learning the visual-semantic joint embedding. We propose a two-stage approach for the task that can augment a typical supervised pair-wise ranking loss based formulation with weakly-annotated web images to learn a more robust visual-semantic embedding. Experiments on two standard benchmark datasets demonstrate that our method achieves a significant performance gain in image-text retrieval compared to state-of-the-art approaches.

📄 PDF Abstract BibTeX arXiv:1808.07793

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage-text RetrievalRetrievalText Retrieval

Similar Papers 제목 키워드 기반

Webly Supervised Joint Embedding for Cross-Modal lmage-Text Retrieval

2018-10-01 · Proceedings of the 26th ACM international conference on Multimedia·October 2018 2018 10 · Niluthpol Chowdhury Mithun, Rameswar Panda, Vagelis Papalexakis, Amit K. Roy-Chowdhury

Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great promise by learning deep representations a…

Cross-Modal RetrievalImage-text RetrievalRetrievalText Retrieval

Omni-sourced Webly-supervised Learning for Video Recognition

2020-03-29 · ECCV 2020 8 · Haodong Duan, Yue Zhao, Yuanjun Xiong, Wentao Liu 외

We introduce OmniSource, a novel framework for leveraging web data to train video recognition models. OmniSource overcomes the barriers between data formats, such as images, short videos, and long untrimmed videos for we…

Action ClassificationAction RecognitionVideo Recognition

Webly Supervised Knowledge Embedding Model for Visual Reasoning

2020-06-01 · CVPR 2020 6 · Wenbo Zheng, Lan Yan, Chao Gou, Fei-Yue Wang

Visual reasoning between visual image and natural language description is a long-standing challenge in computer vision. While recent approaches offer a great promise by compositionality or relational computing, most of t…

modelRepresentation LearningVisual Reasoning

Webly Supervised Fine-Grained Recognition: Benchmark Datasets and An Approach

2021-08-05 · ICCV 2021 10 · Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Yongshun Zhang 외

Learning from the web can ease the extreme dependence of deep learning on large-scale manually labeled datasets. Especially for fine-grained recognition, which targets at distinguishing subordinate categories, it will si…

Benchmarking

Image to Video Domain Adaptation Using Web Supervision

2019-08-05 · Andrew Kae, Yale Song

Training deep neural networks typically requires large amounts of labeled data which may be scarce or expensive to obtain for a particular target domain. As an alternative, we can leverage webly-supervised data (i.e. res…

Domain Adaptation