paper-with-me

Papers

Enhance Modality Robustness in Text-Centric Multimodal Alignment with Adversarial Prompting

2024-08-19 · Yun-Da Tsai, Ting-Yu Yen, Keng-Te Liao, Shou-De Lin

Converting different modalities into generalized text, which then serves as input prompts for large language models (LLMs), is a common approach for aligning multimodal models, particularly when pairwise data is limited. Text-centric alignment method leverages the unique properties of text as a modality space, transforming diverse inputs into a unified textual representation, thereby enabling downstream models to effectively interpret various modal inputs. This study evaluates the quality and robustness of multimodal representations in the face of noise imperfections, dynamic input order permutations, and missing modalities, revealing that current text-centric alignment methods can compromise downstream robustness. To address this issue, we propose a new text-centric adversarial training approach that significantly enhances robustness compared to traditional robust training methods and pre-trained multimodal foundation models. Our findings underscore the potential of this approach to improve the robustness and adaptability of multimodal representations, offering a promising solution for dynamic and real-world applications.

📄 PDF Abstract BibTeX arXiv:2408.09798

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhance the Robustness of Text-Centric Multimodal Alignments

2024-07-06 · Ting-Yu Yen, Yun-Da Tsai, Keng-Te Liao, Shou-De Lin

Converting different modalities into general text, serving as input prompts for large language models (LLMs), is a common method to align multimodal models when there is limited pairwise data. This text-centric approach …

Text-centric Alignment for Multi-Modality Learning

2024-02-12 · Yun-Da Tsai, Ting-Yu Yen, Pei-Fu Guo, Zhe-Yan Li 외

This research paper addresses the challenge of modality mismatch in multimodal learning, where the modalities available during inference differ from those available at training. We propose the Text-centric Alignment for …

In-Context Learning

Knowledge Distillation for Multimodal Egocentric Action Recognition Robust to Missing Modalities

2025-04-11 · Maria Santos-Villafranca, Dustin Carrión-Ojeda, Alejandro Perez-Yus, Jesus Bermudez-Cameo 외

Action recognition is an essential task in egocentric vision due to its wide range of applications across many fields. While deep learning methods have been proposed to address this task, most rely on a single modality, …

Action RecognitionKnowledge Distillation

With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition

2021-11-01 · Evangelos Kazakos, Jaesung Huh, Arsha Nagrani, Andrew Zisserman 외

In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order to improve recognition performance. To in…

Action RecognitionLanguage ModelingLanguage Modelling

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning

2025-07-17 · Suorong Yang, Peijia Li, Yujie Liu, Zhiming Xu 외

Modern deep models are trained on large real-world datasets, where data quality varies and redundancy is common. Data-centric approaches such as dataset pruning have shown promise in improving training efficiency and mod…