paper-with-me

Papers

Efficient Multi-Instance Generation with Janus-Pro-Dirven Prompt Parsing

2025-03-27 · Fan Qi, Yu Duan, Changsheng Xu

Recent advances in text-guided diffusion models have revolutionized conditional image generation, yet they struggle to synthesize complex scenes with multiple objects due to imprecise spatial grounding and limited scalability. We address these challenges through two key modules: 1) Janus-Pro-driven Prompt Parsing, a prompt-layout parsing module that bridges text understanding and layout generation via a compact 1B-parameter architecture, and 2) MIGLoRA, a parameter-efficient plug-in integrating Low-Rank Adaptation (LoRA) into UNet (SD1.5) and DiT (SD3) backbones. MIGLoRA is capable of preserving the base model's parameters and ensuring plug-and-play adaptability, minimizing architectural intrusion while enabling efficient fine-tuning. To support a comprehensive evaluation, we create DescripBox and DescripBox-1024, benchmarks that span diverse scenes and resolutions. The proposed method achieves state-of-the-art performance on COCO and LVIS benchmarks while maintaining parameter efficiency, demonstrating superior layout fidelity and scalability for open-world synthesis.

📄 PDF Abstract BibTeX arXiv:2503.21069

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Image GenerationImage GenerationLayout Generation

Methods 이 논문이 사용한 방법론

BASE 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

2024-10-17 · CVPR 2025 1 · Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma 외

In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due …

Visual Question Answering

Viewpoint Consistency in 3D Generation via Attention and CLIP Guidance

2024-12-03 · Qing Zhang, Zehao Chen, Jinguang Tong, Jing Zhang 외

Despite recent advances in text-to-3D generation techniques, current methods often suffer from geometric inconsistencies, commonly referred to as the Janus Problem. This paper identifies the root cause of the Janus Probl…

3D GenerationText to 3D

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

2025-01-29 · Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan 외

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model s…

Image GenerationInstruction FollowingText to Image Generation+2

Taming Mode Collapse in Score Distillation for Text-to-3D Generation

2023-12-31 · CVPR 2024 1 · Peihao Wang, Dejia Xu, Zhiwen Fan, Dilin Wang 외

Despite the remarkable performance of score distillation in text-to-3D generation, such techniques notoriously suffer from view inconsistency issues, also known as "Janus" artifact, where the generated objects fake each …

3D GenerationPrompt EngineeringText to 3D

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

2025-06-22 · Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen 외

Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabili…

GPUImage GenerationLanguage ModelingLanguage Modelling+4