paper-with-me

Papers

Improved Baselines for Data-efficient Perceptual Augmentation of LLMs

2024-03-20 · Théophane Vallaeys, Mustafa Shukor, Matthieu Cord, Jakob Verbeek

The abilities of large language models (LLMs) have recently progressed to unprecedented levels, paving the way to novel applications in a wide variety of areas. In computer vision, LLMs can be used to prime vision-language tasks such image captioning and visual question answering when coupled with pre-trained vision backbones. While different approaches have been explored to interface LLMs with ``perceptual backbones'' that process, e.g., visual or audio data, they are often explored for different tasks, different datasets, and using different perceptual backbones and language models, hindering direct comparison of the interfacing mechanisms. To remedy this lack of comparability between methods, we present an extensive experimental evaluation of different interfacing mechanisms, across multiple tasks (including image, video, and audio captioning as well as visual question answering), datasets and backbones, paying special attention to low-data settings. We find improved performance using existing mechanisms over state-of-the-art results, and identify a new interfacing mechanism that yields (near) optimal results across different tasks, while obtaining a 4x reduction in training time.

📄 PDF Abstract BibTeX arXiv:2403.13499

Code (0)

등록된 구현이 없습니다.

Tasks

Audio captioningImage CaptioningQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

Is Robustness Robust? On the interaction between augmentations and corruptions

2021-01-01 · Eric Mintun, Alexander Kirillov, Saining Xie

Invariance to a broad array of image corruptions, such as warping, noise, or color shifts, is an important aspect of building robust models in computer vision. Recently, several new data augmentations have been proposed …

SkinGenBench: Generative Model and Preprocessing Effects for Synthetic Dermoscopic Augmentation in Melanoma Diagnosis

2025-12-19 · N. A. Adarsh Pritam, Jeba Shiney O, Sanyam Jain arxiv

This work introduces SkinGenBench, a systematic biomedical imaging benchmark that investigates how preprocessing complexity interacts with generative model choice for synthetic dermoscopic image augmentation and downstre…

Image AugmentationData Augmentation

On Interaction Between Augmentations and Corruptions in Natural Corruption Robustness

2021-02-22 · NeurIPS 2021 12 · Eric Mintun, Alexander Kirillov, Saining Xie

Invariance to a broad array of image corruptions, such as warping, noise, or color shifts, is an important aspect of building robust models in computer vision. Recently, several new data augmentations have been proposed …

DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers

2023-06-15 · Chandramouli Sastry, Sri Harsha Dumpala, Sageev Oore

We introduce DiffAug, a simple and efficient diffusion-based augmentation technique to train image classifiers for the crucial yet challenging goal of improved classifier robustness. Applying DiffAug to a given example c…

DenoisingImage GenerationOut-of-Distribution Detection

The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs

2025-12-17 · Tejas Anvekar, Fenil Bardoliya, Pavan K. Turaga, Chitta Baral 외 arxiv

Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language componen…

Visual GroundingImage Matching