paper-with-me

Papers

NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

2025-08-14 · Shanyuan Liu, Jian Zhu, Junda Lu, Yue Gong, Liuzhuozheng Li, Bo Cheng, Yuhang Ma, Liebucha Wu, Xiaoyu Wu, Dawei Leng, Yuhui Yin arxiv

Diffusion Transformers (DiTs) have demonstrated exceptional capabilities in text-to-image synthesis. However, in the domain of controllable text-to-image generation using DiTs, most existing methods still rely on the ControlNet paradigm originally designed for UNet-based diffusion models. This paradigm introduces significant parameter overhead and increased computational costs. To address these challenges, we propose the Nano Control Diffusion Transformer (NanoControl), which employs Flux as the backbone network. Our model achieves state-of-the-art controllable text-to-image generation performance while incurring only a 0.024\% increase in parameter count and a 0.029\% increase in GFLOPs, thus enabling highly efficient controllable generation. Specifically, rather than duplicating the DiT backbone for control, we design a LoRA-style (low-rank adaptation) control module that directly learns control signals from raw conditioning inputs. Furthermore, we introduce a KV-Context Augmentation mechanism that integrates condition-specific key-value information into the backbone in a simple yet highly effective manner, facilitating deep fusion of conditional features. Extensive benchmark experiments demonstrate that NanoControl significantly reduces computational overhead compared to conventional control approaches, while maintaining superior generation quality and achieving improved controllability.

📄 PDF Abstract BibTeX arXiv:2508.10424

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Beyond Inpainting: Unleash 3D Understanding for Precise Camera-Controlled Video Generation

2026-01-15 · Dong-Yu Chen, Yixin Guo, Shuojin Yang, Tai-Jiang Mu 외 arxiv

Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The ma…

Video Generation

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

2025-12-12 · Han Lin, Xichen Pan, Ziqi Huang, Ji Hou 외 arxiv

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are…

Text-to-Image GenerationVideo Generation

AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models

2025-08-04 · Die Chen, Zhongjie Duan, Zhiwen Li, Cen Chen 외 arxiv

Diffusion models have recently become the dominant paradigm for image generation, yet existing systems struggle to interpret and follow numeric instructions for adjusting semantic attributes. In real-world creative scena…

Reinforcement LearningContinuous ControlImage Generation

FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion

2025-05-21 · Kazuaki Mishima, Antoni Bigata Casademunt, Stavros Petridis, Maja Pantic 외

Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, expression, and emotion. While recent advances in image generation have ena…

AttributeDiversityFace GenerationImage Generation

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

2026-04-09 · Ziqi Cai, Taoyu Yang, Zheng Chang, Si Li 외 arxiv

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only …

Video-to-Video SynthesisVideo Generation