paper-with-me

Papers

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

2026-08-04 · Chunyang Jiang, Pingping Zhang, Yuzhi Zhao, Wenao Ma, Zhijian Hou, Mengyang Wu, Yiyang Cai, Senkang Hu, Sitong Cheng, Chi-Min Chan, Wei Xue, Yike Guo arxiv

Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling models to expand their own training data without external supervision. However, existing MLLM self-augmentation methods are largely text-centric, while image augmentation remains underexplored and typically relies on generic or handcrafted transformations that are weakly aligned with the model's actual incapability. We propose Failure-informed Image Self-Augmentation (\textbf{FISA}), a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases. Our method generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion. Experiments on visual question answering benchmarks show that the proposed method consistently improves performance across both in-distribution and out-of-distribution settings. Further experiments validate the compatibility of FISA with existing textual self-augmentation approaches, the superior data efficiency of the synthesized samples over generic image augmentation baselines, and the practical effectiveness of the proposed filtering strategy.

📄 PDF Abstract BibTeX arXiv:2608.03733

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringImage Augmentation

Similar Papers 제목 키워드 기반

MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science

2025-10-14 · Junkai Zhang, Jingru Gan, Xiaoxuan Wang, Zian Jia 외 arxiv

Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To fill this gap, we introduce MatSciBench, a comprehensive college-level be…

Multimodal Reasoning

GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

2025-02-25 · Theodor Lundqvist, Ludvig Delvret

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and neg…

Representation Learning

MIRROR: Learning from the Other View for Multi-Modal Reasoning

2026-07-23 · Wen Ye, Yuxiao Qu, Aviral Kumar, Xuezhe Ma arxiv

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined di…

Reinforcement LearningMultimodal ReasoningVisual Reasoning

Labels or Input? Rethinking Augmentation in Multimodal Hate Detection

2025-08-15 · Sahajpreet Singh, Kokil Jaidka, Subhayan Mukerjee arxiv

Online hate remains a significant societal challenge, especially as multimodal content enables subtle, culturally grounded, and implicit forms of harm. Hateful memes embed hostility through text-image interactions and hu…

Data Augmentation

Self-Supervised Learning of Motion-Informed Latents

2021-09-29 · Raphaël Jean, Pierre-Luc St-Charles, Soren Pirk, Simon Brodeur

Siamese network architectures trained for self-supervised instance recognition can learn powerful visual representations that are useful in various tasks. Many such approaches work by simply maximizing the similarity bet…

Action RecognitionData AugmentationPose EstimationSelf-Supervised Learning