paper-with-me

Papers

ARMADA: Attribute-Based Multimodal Data Augmentation

2024-08-19 · Xiaomeng Jin, Jeonghwan Kim, Yu Zhou, Kuan-Hao Huang, Te-Lin Wu, Nanyun Peng, Heng Ji

In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal data augmentation frameworks propose ways to augment image-text pairs, they either suffer from semantic inconsistency between texts and images, or generate unrealistic images, causing knowledge gap with real world examples. To address these issues, we propose Attribute-based Multimodal Data Augmentation (ARMADA), a novel multimodal data augmentation method via knowledge-guided manipulation of visual attributes of the mentioned entities. Specifically, we extract entities and their visual attributes from the original text data, then search for alternative values for the visual attributes under the guidance of knowledge bases (KBs) and large language models (LLMs). We then utilize an image-editing model to edit the images with the extracted attributes. ARMADA is a novel multimodal data generation framework that: (i) extracts knowledge-grounded attributes from symbolic KBs for semantically consistent yet distinctive image-text pair generation, (ii) generates visually similar images of disparate categories using neighboring entities in the KB hierarchy, and (iii) uses the commonsense knowledge of LLMs to modulate auxiliary visual attributes such as backgrounds for more robust representation of original entities. Our empirical results over four downstream tasks demonstrate the efficacy of our framework to produce high-quality data and enhance the model performance. This also highlights the need to leverage external knowledge proxies for enhanced interpretability and real-world grounding.

📄 PDF Abstract BibTeX arXiv:2408.10086

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeData Augmentation

Similar Papers 제목 키워드 기반

From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers

2026-03-11 · Ayan Sengupta, Shantanu Dixit, Md Shad Akhtar, Tanmoy Chakraborty arxiv

Knowledge distillation (KD) methods are pivotal in compressing large pre-trained language models into smaller models, ensuring computational efficiency without significantly dropping performance. Traditional KD technique…

Natural Language UnderstandingComputational EfficiencyKnowledge Distillation

ARMADA: Autonomous Online Failure Detection and Human Shared Control Empower Scalable Real-world Deployment and Adaptation

2025-10-02 · Wenye Yu, Jun Lv, Zixi Ying, Yang Jin 외 arxiv

Imitation learning has shown promise in learning from large-scale real-world datasets. However, pretrained policies usually perform poorly without sufficient in-domain data. Besides, human-collected demonstrations entail…

Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks

2025-02-25 · Roger Waleffe, Devesh Sarda, Jason Mohoney, Emmanouil-Vasileios Vlatakis-Gkaragkounis 외

We study distributed training of Graph Neural Networks (GNNs) on billion-scale graphs that are partitioned across machines. Efficient training in this setting relies on min-edge-cut partitioning algorithms, which minimiz…

NARMADA: Need and Available Resource Managing Assistant for Disasters and Adversities

2020-05-27 · WS 2020 7 · Kaustubh Hiware, Ritam Dutt, Sayan Sinha, Sohan Patro 외

Although a lot of research has been done on utilising Online Social Media during disasters, there exists no system for a specific task that is critical in a post-disaster scenario -- identifying resource-needs and resour…

Information RetrievalManagementRetrieval

Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning

2023-06-19 · Shivaen Ramshetty, Gaurav Verma, Srijan Kumar

The robustness of multimodal deep learning models to realistic changes in the input text is critical for their applicability to important tasks such as text-to-image retrieval and cross-modal entailment. To measure robus…

AttributeImage RetrievalMultimodal Deep LearningRetrieval