paper-with-me

홈 › Papers

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

2026-07-28 · Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu arxiv

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unified multimodal model built with low demand on computation and data. Instead of aligning modalities from scratch, Argus-Unified effectively leverages pretrained vision-language models (VLMs) that provide strong multimodal priors. Specifically, we introduce hybrid visual tokens that preserve continuous tokens for understanding while learning discrete tokens for generation from a frozen unified vision encoder. Our training pipeline includes two stages: the first stage learns a quantizer and image decoder on top of the frozen vision encoder, the second stage trains the LLM initialized from a pretrained VLM for the unified multimodal modeling. Using by far the least amount of data (15.6M) and the lowest cost (~$2,000), we demonstrate that unified multimodal models can be trained economically while achieving strong performance in both understanding and generation. Notably, our model attains state-of-the-art multimodal understanding on GQA, POPE, and VQAv2, and competitive generation quality compared to models with dedicated vision encoders (e.g., Janus, Janus-Pro), all at ~10x lower cost and with ~5x less data. We envision Argus-Unified as a useful baseline that lowers the development barrier for unified models.

📄 PDF Abstract BibTeX arXiv:2607.25527

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Argus: A Compact and Versatile Foundation Model for Vision

2025-01-01 · CVPR 2025 1 · Weiming Zhuang, Chen Chen, Zhizhong Li, Sina Sajadmanesh 외

While existing vision and multi-modal foundation models can handle multiple computer vision tasks, they often suffer from significant limitations, including huge demand for data and computational resources during tra…

Smart obervation method with wide field small aperture telescopes for real time transient detection

2020-11-20 · Peng Jia, Qiang Liu, Yongyang Sun, Yitian Zheng 외

Wide field small aperture telescopes (WFSATs) are commonly used for fast sky survey. Telescope arrays composed by several WFSATs are capable to scan sky several times per night. Huge amount of data would be obtained by t…

Ensemble Learning

What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media

2026-06-05 · Zifan Peng, Yini Huang, Aiwen Lu, Qiming Ye 외 arxiv

Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulative and cross-post: cues that appear harmless in isolation may jointly e…

Empirical Recipes for Efficient and Compact Vision-Language Models

2026-03-17 · Jiabo Huang, Zhizhong Li, Sina Sajadmanesh, Weiming Zhuang 외 arxiv

Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedups their smaller parameter counts sugges…

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

2026-06-10 · Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong 외 arxiv

Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and c…

Video Generation