paper-with-me

홈 › Papers

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

2026-05-20 · Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, Chong Luo, Jianmin Bao, Ji Li, Lei Shi, Qinhong Yang, Xiuyu Wu, Xuelu Feng, Yan Lu, Yanchen Dong, Yitong Wang, Yunuo Chen arxiv

We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters across various benchmarks, while requiring significantly less training compute. For example, Lens requires only about 19.3% of the training compute used by Z-Image. The training efficiency of Lens stems from two key strategies beyond its compact model size. First, we maximize data information density per training batch by (i) training on Lens-800M, a dataset of 800M densely captioned image-text pairs whose captions are generated by GPT-4.1 and contain approximately 109 words on average, providing richer semantic supervision than conventional short captions, and (ii) constructing each batch from images with multiple resolutions and diverse aspect ratios, thereby enlarging the effective visual coverage of each optimization step. Second, we improve convergence speed through careful architectural choices, including adopting a semantic VAE that provides better latent representations and employing a strong language encoder that accelerates optimization while enabling multilingual generalization from English-only training data. After pre-training, we apply RL with taxonomy-driven prompts (Lens-RL-8K) and structured reward rubrics to suppress artifacts and improve visual quality, a reasoner module with training-free system prompt search to better align user requests with the model, and distillation-based acceleration for 4-step inference. Through efficient training and systematic optimization, Lens generalizes to arbitrary aspect ratios from 1:2 to 2:1 and resolutions up to 1440^2, and supports prompts in several commonly used languages. Thanks to its compact size, Lens generates a 1024^2 image in 3.15 seconds on a single NVIDIA H100 GPU, while its distilled turbo version performs 4-step generation in 0.84 seconds.

📄 PDF Abstract BibTeX arXiv:2605.21573

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models

2023-05-28 · Zhiwei Jia, Pradyumna Narayana, Arjun R. Akula, Garima Pruthi 외

Image ad understanding is a crucial task with wide real-world applications. Although highly challenging with the involvement of diverse atypical scenes, real-world entities, and reasoning over scene-texts, how to interpr…

Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes

2025-10-26 · Guanyu Yao, Qiucheng Wu, Yang Zhang, Zhaowen Wang 외 arxiv

Multimodal large language models (MLLMs) have demonstrated strong capabilities on vision-and-language tasks. However, recent findings reveal an imbalance in their reasoning capabilities across visual and textual modaliti…

Multimodal ReasoningVisual Reasoning

Understanding Machine Learning Paradigms through the Lens of Statistical Thermodynamics: A tutorial

2024-11-24 · Star, Liu

This tutorial investigates the convergence of statistical mechanics and learning theory, elucidating the potential enhancements in machine learning methodologies through the integration of foundational principles from ph…

Learning TheoryVariational Inference

TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

2025-12-16 · Jun Zhang, Teng Wang, Yuying Ge, Yixiao Ge 외 arxiv

This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in video understanding. While multimodal l…

Reinforcement Learning

In Dialogue with Intelligence: Rethinking Large Language Models as Collective Knowledge

2025-05-28 · Eleni Vasilaki

Large Language Models (LLMs) are typically analysed through architectural, behavioural, or training-data lenses. This article offers a theoretical and experiential re-framing: LLMs as dynamic instantiations of Collective…