paper-with-me

Papers

MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices

2023-11-28 · Yang Zhao, Yanwu Xu, Zhisheng Xiao, HaoLin Jia, Tingbo Hou

The deployment of large-scale text-to-image diffusion models on mobile devices is impeded by their substantial model size and slow inference speed. In this paper, we propose \textbf{MobileDiffusion}, a highly efficient text-to-image diffusion model obtained through extensive optimizations in both architecture and sampling techniques. We conduct a comprehensive examination of model architecture design to reduce redundancy, enhance computational efficiency, and minimize model's parameter count, while preserving image generation quality. Additionally, we employ distillation and diffusion-GAN finetuning techniques on MobileDiffusion to achieve 8-step and 1-step inference respectively. Empirical studies, conducted both quantitatively and qualitatively, demonstrate the effectiveness of our proposed techniques. MobileDiffusion achieves a remarkable \textbf{sub-second} inference speed for generating a $512\times512$ image on mobile devices, establishing a new state of the art.

📄 PDF Abstract BibTeX arXiv:2311.16567

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyImage GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

InstantIR: Blind Image Restoration with Instant Generative Reference

2024-10-09 · Jen-Yuan Huang, Haofan Wang, Qixun Wang, Xu Bai 외

Handling test-time unknown degradation is the major challenge in Blind Image Restoration (BIR), necessitating high model generalization. An effective strategy is to incorporate prior knowledge, either from human input or…

Image Restoration

InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

2023-04-06 · CVPR 2024 1 · Jing Shi, Wei Xiong, Zhe Lin, Hyun Joon Jung

Recent advances in personalized image generation allow a pre-trained text-to-image model to learn a new concept from a set of images. However, existing personalization approaches usually require heavy test-time finetunin…

Diffusion PersonalizationDiffusion Personalization Tuning FreeImage GenerationPersonalized Image Generation+2

InstantDrag: Improving Interactivity in Drag-based Image Editing

2024-09-13 · Joonghyuk Shin, Daehyeon Choi, Jaesik Park

Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to…

Image GenerationMotion GenerationOptical Flow Estimation

Mobile Phone Based Vehicle License Plate Recognition for Road Policing

2015-04-07 · Lajish V. L., Sunil Kumar Kopparapu

Identity of a vehicle is done through the vehicle license plate by traffic police in general. Au- tomatic vehicle license plate recognition has several applications in intelligent traffic management systems. The security…

License Plate RecognitionManagement

InstantID: Zero-shot Identity-Preserving Generation in Seconds

2024-01-15 · Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin 외

There has been significant progress in personalized image synthesis with methods such as Textual Inversion, DreamBooth, and LoRA. Yet, their real-world applicability is hindered by high storage demands, lengthy fine-tuni…

Diffusion PersonalizationDiffusion Personalization Tuning FreeImage Generation