paper-with-me

Papers

Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

2018-09-21 · Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent advances in learning representations from cross-modal self-supervision. The main contributions of this paper are as follows: (1) we propose a new learning strategy where the embeddings are learnt via a multi-way matching problem, as opposed to a binary classification (matching or non-matching) problem as proposed by recent papers; (2) we demonstrate that performance of this method far exceeds the existing baselines on the synchronization task; (3) we use the learnt embeddings for visual speech recognition in self-supervision, and show that the performance matches the representations learnt end-to-end in a fully-supervised manner.

📄 PDF Abstract BibTeX arXiv:1809.08001

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationCross-Modal RetrievalRetrievalspeech-recognitionSpeech RecognitionVideo SynchronizationVisual Speech Recognition

Similar Papers 제목 키워드 기반

Improved Probabilistic Image-Text Representations

2023-05-29 · Sanghyuk Chun

Image-Text Matching (ITM) task, a fundamental vision-language (VL) task, suffers from the inherent ambiguity arising from multiplicity and imperfect annotations. Deterministic functions are not sufficiently powerful to c…

Data AugmentationImage-text matchingText Matchingzero-shot-classification+1

Deep Cross-Modal Projection Learning for Image-Text Matching

2018-09-01 · ECCV 2018 9 · Ying Zhang, Huchuan Lu

The key point of image-text matching is how to accurately measure the similarity between visual and textual inputs. Despite the great progress of associating the deep cross-modal embeddings with the bi-directional rankin…

Cross-Modal RetrievalImage-text matchingText based Person RetrievalText Matching

Rethinking Benchmarks for Cross-modal Image-text Retrieval

2023-04-21 · Weijing Chen, Linli Yao, Qin Jin

Image-text retrieval, as a fundamental and important branch of information retrieval, has attracted extensive research attentions. The main challenge of this task is cross-modal semantic understanding and matching. Some …

Cross-Modal RetrievalImage-text RetrievalImage-to-Text RetrievalInformation Retrieval+2

Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning

2025-01-01 · CVPR 2025 1 · Jeong Ryong Lee, Yejee Shin, Geonhui Son, Dosik Hwang

The modality gap between vision and text embeddings in CLIP presents a significant challenge for zero-shot image captioning, limiting effective cross-modal representation. Traditional approaches, such as noise inject…

cross-modal alignmentDenoisingImage Captioning

Uncertainty-based Cross-Modal Retrieval with Probabilistic Representations

2022-04-20 · Leila Pishdad, Ran Zhang, Konstantinos G. Derpanis, Allan Jepson 외

Probabilistic embeddings have proven useful for capturing polysemous word meanings, as well as ambiguity in image matching. In this paper, we study the advantages of probabilistic embeddings in a cross-modal setting (i.e…

Cross-Modal RetrievalImage RetrievalImage-text matchingImage to text+2