paper-with-me

Papers

Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning

2023-06-19 · Shivaen Ramshetty, Gaurav Verma, Srijan Kumar

The robustness of multimodal deep learning models to realistic changes in the input text is critical for their applicability to important tasks such as text-to-image retrieval and cross-modal entailment. To measure robustness, several existing approaches edit the text data, but do so without leveraging the cross-modal information present in multimodal data. Information from the visual modality, such as color, size, and shape, provide additional attributes that users can include in their inputs. Thus, we propose cross-modal attribute insertions as a realistic perturbation strategy for vision-and-language data that inserts visual attributes of the objects in the image into the corresponding text (e.g., "girl on a chair" to "little girl on a wooden chair"). Our proposed approach for cross-modal attribute insertions is modular, controllable, and task-agnostic. We find that augmenting input text using cross-modal insertions causes state-of-the-art approaches for text-to-image retrieval and cross-modal entailment to perform poorly, resulting in relative drops of 15% in MRR and 20% in $F_1$ score, respectively. Crowd-sourced annotations demonstrate that cross-modal insertions lead to higher quality augmentations for multimodal data than augmentations using text-only data, and are equivalent in quality to original examples. We release the code to encourage robustness evaluations of deep vision-and-language models: https://github.com/claws-lab/multimodal-robustness-xmai.

📄 PDF Abstract BibTeX arXiv:2306.11065

Code (1)

claws-lab/multimodal-robustness-xmai 공식 구현 pytorch

Tasks

AttributeImage RetrievalMultimodal Deep LearningRetrieval

Similar Papers 제목 키워드 기반

Primate specific retrotransposons, SVAs, in the evolution of networks that alter brain function

2016-02-27

The hominid-specific non-LTR retrotransposon termed SINE VNTR Alu (SVA) is the youngest of the transposable elements in the human genome. The propagation of the most ancient SVA type A took place about thirteen millions …

MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models

2025-02-12 · Peng-Fei Zhang, Guangdong Bai, Zi Huang

Current adversarial attacks for evaluating the robustness of vision-language pre-trained (VLP) models in multi-modal tasks suffer from limited transferability, where attacks crafted for a specific model often struggle to…

Adversarial Attack

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

2025-07-09 · Ananya Raval, Aravind Narayanan, Vahid Reza Khazaie, Shaina Raza

Large Multimodal Models (LMMs) are typically trained on vast corpora of image-text data but are often limited in linguistic coverage, leading to biased and unfair outputs across languages. While prior work has explored m…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Evaluating Attribute Confusion in Fashion Text-to-Image Generation

2025-07-09 · Ziyue Liu, Federico Girella, Yiming Wang, Davide Talon

Despite the rapid advances in Text-to-Image (T2I) generation models, their evaluation remains challenging in domains like fashion, involving complex compositional generation. Recent automated T2I evaluation methods lever…

Attributecross-modal alignmentImage GenerationQuestion Answering+5

MMPE: A Multi-Modal Interface for Post-Editing Machine Translation

2020-07-01 · ACL 2020 6 · Nico Herbig, Tim D{\"u}wel, Santanu Pal, Kalliopi Meladaki 외

Current advances in machine translation (MT) increase the need for translators to switch from traditional translation to post-editing (PE) of machine-translated text, a process that saves time and reduces errors. This af…

Machine TranslationTranslation