paper-with-me

Papers

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

2025-10-22 · Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, Zhe Gan arxiv

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-400K, a comprehensive 400K-image dataset for instruction-based image editing. Our dataset is constructed by leveraging Nano-Banana to generate diverse edit pairs from real photographs in the OpenImages collection. What distinguishes Pico-Banana-400K from previous synthetic datasets is our systematic approach to quality and diversity. We employ a fine-grained image editing taxonomy to ensure comprehensive coverage of edit types while maintaining precise content preservation and instruction faithfulness through MLLM-based quality scoring and careful curation. Beyond single turn editing, Pico-Banana-400K enables research into complex editing scenarios. The dataset includes three specialized subsets: (1) a 72K-example multi-turn collection for studying sequential editing, reasoning, and planning across consecutive modifications; (2) a 56K-example preference subset for alignment research and reward model training; and (3) paired long-short editing instructions for developing instruction rewriting and summarization capabilities. By providing this large-scale, high-quality, and task-rich resource, Pico-Banana-400K establishes a robust foundation for training and benchmarking the next generation of text-guided image editing models.

📄 PDF Abstract BibTeX arXiv:2510.19808

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation

2025-11-28 · Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa 외 arxiv

Recent text-to-image generation models have acquired the ability of multi-reference generation and editing; that is, to inherit the appearance of subjects from multiple reference images and re-render them in new contexts…

Text-to-Image Generation

Med-Banana: Learning Quality-Controlled Medical Image Editing from Success-and-Failure Trajectories

2025-11-02 · Zhihui Chen, Qingyuan Lei, Kai He, Yanrui Du 외 arxiv

Text-guided medical image editing must satisfy the requested pathology while preserving anatomy, modality-specific appearance, and clinical plausibility. However, existing datasets largely supervise editors with final ac…

Image Editing

Personalized Vision via Visual In-Context Learning

2025-09-29 · Yuxin Jiang, Yuchao Gu, Yiren Song, Ivor Tsang 외 arxiv

Modern vision models, trained on large-scale annotated datasets, excel at predefined tasks but struggle with personalized vision -- tasks defined at test time by users with customized objects or novel objectives. Existin…

Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets

2025-12-17 · Jialong Zuo, Haoyou Deng, Hanyu Zhou, Jiaxin Zhu 외 arxiv

The rapid evolution of text-to-image generation models has revolutionized visual content creation. While commercial products like Nano Banana Pro have garnered significant attention, their potential as generalist solvers…

Text-to-Image Generation

Computing High Accuracy Power Spectra with Pico

2007-12-02 · William A. Fendt, Benjamin D. Wandelt

This paper presents the second release of Pico (Parameters for the Impatient COsmologist). Pico is a general purpose machine learning code which we have applied to computing the CMB power spectra and the WMAP likelihood.…

Distributed ComputingPICOVocal Bursts Intensity Prediction