paper-with-me

홈 › Papers

Probing Representations Learned by Multimodal Recurrent and Transformer Models

2019-08-29 · Jindřich Libovický, Pranava Madhyastha

Recent literature shows that large-scale language modeling provides excellent reusable sentence representations with both recurrent and self-attentive architectures. However, there has been less clarity on the commonalities and differences in the representational properties induced by the two architectures. It also has been shown that visual information serves as one of the means for grounding sentence representations. In this paper, we present a meta-study assessing the representational quality of models where the training signal is obtained from different modalities, in particular, language modeling, image features prediction, and both textual and multimodal machine translation. We evaluate textual and visual features of sentence representations obtained using predominant approaches on image retrieval and semantic textual similarity. Our experiments reveal that on moderate-sized datasets, a sentence counterpart in a target language or visual modality provides much stronger training signal for sentence representation than language modeling. Importantly, we observe that while the Transformer models achieve superior machine translation quality, representations from the recurrent neural network based models perform significantly better over tasks focused on semantic relevance.

📄 PDF Abstract BibTeX arXiv:1908.11125

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalLanguage ModelingLanguage ModellingMachine TranslationMultimodal Machine TranslationRetrievalSemantic Textual SimilaritySentenceTranslation

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking

2024-01-29 · Ivana Beňová, Jana Košecká, Michal Gregor, Martin Tamajka 외

The dominant probing approaches rely on the zero-shot performance of image-text matching tasks to gain a finer-grained understanding of the representations learned by recent multimodal image-language transformer models. …

Image-text matchingText Matching

Probing Cross-Modal Representations in Multi-Step Relational Reasoning

2021-08-01 · ACL (RepL4NLP) 2021 8 · Iuliia Parfenova, Desmond Elliott, Raquel Fernández, Sandro Pezzelle

We investigate the representations learned by vision and language models in tasks that require relational reasoning. Focusing on the problem of assessing the relative size of objects in abstract visual contexts, we analy…

DiagnosticRelational Reasoning

Multi-layer attentive probing improves transfer of audio representations for bioacoustics

2026-05-11 · Marius Miron, David Robinson, Masato Hagiwara, Titouan Parcollet 외 arxiv

Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-c…

Representation Learning

Probing Graph Representations

2023-03-07 · Mohammad Sadegh Akhondzadeh, Vijay Lingam, Aleksandar Bojchevski

Today we have a good theoretical understanding of the representational power of Graph Neural Networks (GNNs). For example, their limitations have been characterized in relation to a hierarchy of Weisfeiler-Lehman (WL) is…

Diagnostic

Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models

2022-12-01 · Zhuowan Li, Cihang Xie, Benjamin Van Durme, Alan Yuille

Despite the impressive advancements achieved through vision-and-language pretraining, it remains unclear whether this joint learning paradigm can help understand each individual modality. In this work, we conduct a compa…

AttributePredictionRepresentation Learning