paper-with-me

Papers

SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

2024-05-07 · Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge, Ying Shan

In this technical report, we introduce SEED-Data-Edit: a unique hybrid dataset for instruction-guided image editing, which aims to facilitate image manipulation using open-form language. SEED-Data-Edit is composed of three distinct types of data: (1) High-quality editing data produced by an automated pipeline, ensuring a substantial volume of diverse image editing pairs. (2) Real-world scenario data collected from the internet, which captures the intricacies of user intentions for promoting the practical application of image editing in the real world. (3) High-precision multi-turn editing data annotated by humans, which involves multiple rounds of edits for simulating iterative editing processes. The combination of these diverse data sources makes SEED-Data-Edit a comprehensive and versatile dataset for training language-guided image editing model. We fine-tune a pretrained Multimodal Large Language Model (MLLM) that unifies comprehension and generation with SEED-Data-Edit. The instruction tuned model demonstrates promising results, indicating the potential and effectiveness of SEED-Data-Edit in advancing the field of instructional image editing. The datasets are released in https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit.

📄 PDF Abstract BibTeX arXiv:2405.04007

Code (1)

ailab-cvc/seed-x 공식 구현 pytorch

Tasks

Image ManipulationLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

Similar Papers 제목 키워드 기반

Step-Audio-EditX Technical Report

2025-11-05 · Chao Yan, Boyong Wu, Peng Yang, Pengfei Tan 외 arxiv

We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguistics alongside robust zero-shot text-to-sp…

SeedEdit 3.0: Fast and High-Quality Generative Image Editing

2025-06-05 · Peng Wang, Yichun Shi, Xiaochen Lian, Zhonghua Zhai 외

We introduce SeedEdit 3.0, in companion with our T2I model Seedream 3.0, which significantly improves over our previous SeedEdit versions in both aspects of edit instruction following and image content (e.g., ID/IP) pres…

Instruction Following

Technical Report: Hybrid Autonomous Intersection Management

2022-04-16 · Aaron Parks-Young, Guni Sharon

This document provides technical details regarding the Hybrid-AIM simulator that was used in Sharon and Stone (2017) and Parks-Young and Sharon (2022).

Management

Seed1.5-VL Technical Report

2025-05-11 · Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng 외

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (…

Mixture-of-ExpertsMultimodal ReasoningVideo Question AnsweringVideo Understanding

Reward-Modulated Local Learning in Spiking Encoders: Controlled Benchmarks with STDP and Hybrid Rate Readouts

2026-02-28 · Debjyoti Chakraborty arxiv

This paper presents a controlled empirical study of biologically motivated local learning for handwritten digit recognition. We evaluate an STDP-inspired competitive proxy and a practical hybrid benchmark built on the sa…

Handwritten Digit Recognition