paper-with-me

홈 › Papers

Order embeddings and character-level convolutions for multimodal alignment

2017-06-03 · Jônatas Wehrmann, Anderson Mattjie, Rodrigo C. Barros

With the novel and fast advances in the area of deep neural networks, several challenging image-based tasks have been recently approached by researchers in pattern recognition and computer vision. In this paper, we address one of these tasks, which is to match image content with natural language descriptions, sometimes referred as multimodal content retrieval. Such a task is particularly challenging considering that we must find a semantic correspondence between captions and the respective image, a challenge for both computer vision and natural language processing areas. For such, we propose a novel multimodal approach based solely on convolutional neural networks for aligning images with their captions by directly convolving raw characters. Our proposed character-based textual embeddings allow the replacement of both word-embeddings and recurrent neural networks for text understanding, saving processing time and requiring fewer learnable parameters. Our method is based on the idea of projecting both visual and textual information into a common embedding space. For training such embeddings we optimize a contrastive loss function that is computed to minimize order-violations between images and their respective descriptions. We achieve state-of-the-art performance in the largest and most well-known image-text alignment dataset, namely Microsoft COCO, with a method that is conceptually much simpler and that possesses considerably fewer parameters than current approaches.

📄 PDF Abstract BibTeX arXiv:1706.00999

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalSemantic correspondenceWord Embeddings

Similar Papers 제목 키워드 기반

Bidirectional Retrieval Made Simple

2018-06-01 · CVPR 2018 6 · Jônatas Wehrmann, Rodrigo C. Barros

This paper provides a very simple yet effective character-level architecture for learning bidirectional retrieval models. Aligning multimodal content is particularly challenging considering the difficulty in finding sema…

Image RetrievalRetrievalSemantic correspondencetext-classification+2

Learning distributed sentence vectors with bi-directional 3D convolutions

2020-12-01 · COLING 2020 8 · Bin Liu, Liang Wang, Guosheng Yin

We propose to learn distributed sentence representation using text{'}s visual features as input. Different from the existing methods that render the words or characters of a sentence into images separately, we further fo…

SentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings

Learning semantic sentence representations from visually grounded language without lexical knowledge

2019-03-27 · Danny Merkx, Stefan Frank

Current approaches to learning semantic representations of sentences often use prior word-level knowledge. The current study aims to leverage visual information in order to capture sentence level semantics without the ne…

Grounded language learningLearning Semantic RepresentationsRetrievalSemantic Similarity+5

CAMEL-CLIP: Channel-aware Multimodal Electroencephalography-text Alignment for Generalizable Brain Foundation Models

2026-02-27 · Hanseul Choi, Jinyeong Park, Seongwon Jin, Sungho Park 외 arxiv

Electroencephalography (EEG) foundation models have shown promise for learning generalizable representations, yet they remain sensitive to channel heterogeneity, such as changes in channel composition or ordering. We pro…

Contrastive Learning

Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network

2023-02-26 · Xing Yi Liu, Homayoon Beigi

Punctuation restoration plays an essential role in the post-processing procedure of automatic speech recognition, but model efficiency is a key requirement for this task. To that end, we present EfficientPunct, an ensemb…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Computational EfficiencyPunctuation Restoration+2