paper-with-me

홈 › Papers

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

2023-10-01 · Tianyu Yu, Jinyi Hu, Yuan YAO, Haoye Zhang, Yue Zhao, Chongyi Wang, Shan Wang, Yinxv Pan, Jiao Xue, Dahai Li, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature alignment of visual modules and large language models; the multimodal instruction tuning datasets for human instruction following. (i) For the model architecture, most existing models introduce an external bridge module to connect vision encoders with language models, which needs an additional feature-alignment pre-training. In this work, we discover that compact pre-trained vision language models can inherently serve as ``out-of-the-box'' bridges between vision and language. Based on this, we propose Muffin framework, which directly employs pre-trained vision-language models to act as providers of visual signals. (ii) For the multimodal instruction tuning datasets, existing methods omit the complementary relationship between different datasets and simply mix datasets from different tasks. Instead, we propose UniMM-Chat dataset which explores the complementarities of datasets to generate 1.1M high-quality and diverse multimodal instructions. We merge information describing the same image from diverse datasets and transforms it into more knowledge-intensive conversation data. Experimental results demonstrate the effectiveness of the Muffin framework and UniMM-Chat dataset. Muffin achieves state-of-the-art performance on a wide range of vision-language tasks, significantly surpassing state-of-the-art models like LLaVA and InstructBLIP. Our model and dataset are all accessible at https://github.com/thunlp/muffin.

📄 PDF Abstract BibTeX arXiv:2310.00653

Code (2)

thunlp/muffin 공식 구현 pytorch
rlhf-v/rlhf-v

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters

2024-08-30 · Mouxiang Chen, Lefei Shen, Zhuo Li, Xiaoyun Joy Wang 외

Foundation models have emerged as a promising approach in time series forecasting (TSF). Existing approaches either repurpose large language models (LLMs) or build large-scale time series datasets to develop TSF foundati…

Image ReconstructionTime SeriesTime Series Forecasting

Aligning and Prompting Everything All at Once for Universal Visual Perception

2023-12-04 · CVPR 2024 1 · Yunhang Shen, Chaoyou Fu, Peixian Chen, Mengdan Zhang 외

Vision foundation models have been explored recently to build general-purpose vision systems. However, predominant paradigms, driven by casting instance-level tasks as an object-word alignment, bring heavy cross-modality…

AllObjectobject-detectionObject Detection+5

VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones

2025-08-06 · Lefei Shen, Mouxiang Chen, Xu Liu, Han Fu 외 arxiv

Recent studies have indicated that vision models pre-trained on images can serve as time series foundation models (TSFMs) by reformulating time series forecasting (TSF) as image reconstruction. However, effective cross-m…

Time Series ForecastingImage Reconstruction

TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding

2026-05-06 · Minjie Qiang, Mingming Zhang, Xiaoyi Bao, Xing Fu 외 arxiv

Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approach…

Representation LearningContrastive Learning

FLAVA: A Foundational Language And Vision Alignment Model

2021-12-08 · CVPR 2022 1 · Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 외

State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal…

Image RetrievalImage-to-Text RetrievalVisual ReasoningZero-shot Image Retrieval+2