paper-with-me

홈 › Papers

Pinterest Canvas: Large-Scale Image Generation at Pinterest

2026-03-06 · Yu Wang, Eric Tzeng, Raymond Shiau, Jie Yang, Dmitry Kislyuk, Charles Rosenberg arxiv

While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes them hard to control via prompting or simple inference adaptation alone, rendering them unsuitable for use cases with strict product requirements. In this paper, we introduce Pinterest Canvas, our large-scale image generation system built to support image editing and enhancement use cases at Pinterest. Canvas is first trained on a diverse, multimodal dataset to produce a foundational diffusion model with broad image-editing capabilities. However, rather than relying on one generic model to handle every downstream task, we instead rapidly fine-tune variants of this base model on task-specific datasets, producing specialized models for individual use cases. We describe key components of Canvas and summarize our best practices for dataset curation, training, and inference. We also showcase task-specific variants through case studies on background enhancement and aspect-ratio outpainting, highlighting how we tackle their specific product requirements. Online A/B experiments demonstrate that our enhanced images receive a significant 18.0% and 12.5% engagement lift, respectively, and comparisons with human raters further validate that our models outperform third-party models on these tasks. Finally, we showcase other Canvas variants, including multi-image scene synthesis and image-to-video generation, demonstrating that our approach can generalize to a wide variety of potential downstream tasks.

📄 PDF Abstract BibTeX arXiv:2603.06453

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationImage GenerationImage Editing

Similar Papers 제목 키워드 기반

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

2026-07-06 · Hairui Zhu, Yiying Yang, Tengjin Weng, Ziyu Lu 외 hf

Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositi…

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

2026-07-09 · Haoran Feng, Ruiyang Zhang, Longyi Zhang, Dizhe Zhang 외 arxiv

In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-q…

Style Transfer

PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest

2026-03-03 · Josh Beal, Eric Kim, Jinfeng Rao, Rex Wu 외 arxiv

While multi-modal Visual Language Models (VLMs) have demonstrated significant success across various domains, the integration of VLMs into recommendation and retrieval systems remains a challenge, due to issues like trai…

Representation Learning

Canvas-to-Image: Compositional Image Generation with Multimodal Controls

2025-11-26 · Yusuf Dalva, Guocheng Gordon Qian, Maya Goldenberg, Tsai-Shien Chen 외 arxiv

While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts,…

Text-to-Image GenerationSpatial Reasoning

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

2025-12-12 · Han Lin, Xichen Pan, Ziqi Huang, Ji Hou 외 arxiv

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are…

Text-to-Image GenerationVideo Generation