paper-with-me

홈 › Papers

Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents

2025-02-27 · Zhenyu Liu, Yunxin Li, Baotian Hu, Wenhan Luo, YaoWei Wang, Min Zhang

To improve Multimodal Large Language Models' (MLLMs) ability to process images and complex instructions, researchers predominantly curate large-scale visual instruction tuning datasets, which are either sourced from existing vision tasks or synthetically generated using LLMs and image descriptions. However, they often suffer from critical flaws, including misaligned instruction-image pairs and low-quality images. Such issues hinder training efficiency and limit performance improvements, as models waste resources on noisy or irrelevant data with minimal benefit to overall capability. To address this issue, we propose a \textbf{Vi}sual-Centric \textbf{S}election approach via \textbf{A}gents Collaboration (ViSA), which centers on image quality assessment and image-instruction relevance evaluation. Specifically, our approach consists of 1) an image information quantification method via visual agents collaboration to select images with rich visual information, and 2) a visual-centric instruction quality assessment method to select high-quality instruction data related to high-quality images. Finally, we reorganize 80K instruction data from large open-source datasets. Extensive experiments demonstrate that ViSA outperforms or is comparable to current state-of-the-art models on seven benchmarks, using only 2.5\% of the original data, highlighting the efficiency of our data selection approach. Moreover, we conduct ablation studies to validate the effectiveness of each component of our method. The code is available at https://github.com/HITsz-TMG/ViSA.

📄 PDF Abstract BibTeX arXiv:2502.19917

Code (1)

hitsz-tmg/visa 공식 구현 pytorch

Tasks

Image Quality Assessment

Similar Papers 제목 키워드 기반

Human-centric Image Cropping with Partition-aware and Content-preserving Features

2022-07-21 · Bo Zhang, Li Niu, Xing Zhao, Liqing Zhang

Image cropping aims to find visually appealing crops in an image, which is an important yet challenging task. In this paper, we consider a specific and practical application: human-centric image cropping, which focuses o…

Image Cropping

Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search

2020-10-29 · NeurIPS 2020 12 · Houwen Peng, Hao Du, Hongyuan Yu, Qi Li 외

One-shot weight sharing methods have recently drawn great attention in neural architecture search due to high efficiency and competitive performance. However, weight sharing across models has an inherent deficiency, i.e.…

Neural Architecture Searchobject-detectionObject Detection

Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

2023-05-24 · Geewook Kim, Hodong Lee, Daehee Kim, Haeji Jung 외

Recent advances in Large Language Models (LLMs) have stimulated a surge of research aimed at extending their applications to the visual domain. While these models exhibit promise in generating abstract image captions and…

document understandingImage CaptioningNatural Language UnderstandingQuestion Answering+1

3D Hand Pose Estimation in Everyday Egocentric Images

2023-12-11 · Aditya Prakash, Ruisen Tu, Matthew Chang, Saurabh Gupta

3D hand pose estimation in everyday egocentric images is challenging for several reasons: poor visual signal (occlusion from the object of interaction, low resolution & motion blur), large perspective distortion (hands a…

3D Hand Pose EstimationHand Pose EstimationPose Estimation

ShotCrop$^3$: Cropping Human-Centric Images into Cinematic Triple-Shot Compositions

2026-06-04 · Dehong Kong, Lina Lei, Lingtao Zheng, Chenyang Wu 외 arxiv

Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots from one scene. In practice, multi-shot composition is critical for…