Probabilistic Multimodal Representation Learning
Learning multimodal representations is a requirement for many tasks such as image--caption retrieval. Previous work on this problem has only focused on finding good vector representations without any explicit measure of uncertainty. In this work, we argue and demonstrate that learning multimodal representations as probability distributions can lead to better representations, as well as providing other benefits such as adding a measure of uncertainty to the learned representations. We show that this measure of uncertainty can capture how confident our model is about the representations in the multimodal domain, i.e, how clear it is for the model to retrieve/predict the matching pair. We experiment with similarity metrics that have not been traditionally used for the multimodal retrieval task, and show that the choice of the similarity metric affects the quality of the learned representations.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningRetrievalSimilar Papers 제목 키워드 기반
ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
We introduce ProM3E, a probabilistic masked multimodal embedding model for any-to-any generation of multimodal representations for ecology. ProM3E is based on masked modality reconstruction in the embedding space, learni…
Representation LearningCross-Modal RetrievalContext-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting
Time series forecasting is important for applications spanning energy markets, climate analysis, and traffic management. However, existing methods struggle to effectively integrate exogenous texts and align them with the…
Time SeriesTime Series ForecastingLearning Sequential Latent Variable Models from Multimodal Time Series Data
Sequential modelling of high-dimensional data is an important problem that appears in many domains including model-based reinforcement learning and dynamics identification for control. Latent variable models applied to s…
Model-based Reinforcement LearningTime SeriesTime Series AnalysisProbabilistic Compositional Embeddings for Multimodal Image Retrieval
Existing works in image retrieval often consider retrieving images with one or two query inputs, which do not generalize to multiple queries. In this work, we investigate a more challenging scenario for composing multipl…
Image RetrievalRetrievalA Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images
Several recent studies have shown the benefits of combining language and perception to infer word embeddings. These multimodal approaches either simply combine pre-trained textual and visual representations (e.g. feature…
Coreference ResolutionImage ClassificationQuestion AnsweringRetrieval+3