paper-with-me

Papers

MAGICK: A Large-scale Captioned Dataset from Matting Generated Images using Chroma Keying

2024-01-01 · CVPR 2024 1 · Ryan D. Burgert, Brian L. Price, Jason Kuen, Yijun Li, Michael S. Ryoo

We introduce MAGICK a large-scale dataset of generated objects with high-quality alpha mattes. While image generation methods have produced segmentations they cannot generate alpha mattes with accurate details in hair fur and transparencies. This is likely due to the small size of current alpha matting datasets and the difficulty in obtaining ground-truth alpha. We propose a scalable method for synthesizing images of objects with high-quality alpha that can be used as a ground-truth dataset. A key idea is to generate objects on a single-colored background so chroma keying approaches can be used to extract the alpha. However this faces several challenges including that current text-to-image generation methods cannot create images that can be easily chroma keyed and that chroma keying is an underconstrained problem that generally requires manual intervention for high-quality results. We address this using a combination of generation and alpha extraction methods. Using our method we generate a dataset of 150000 objects with alpha. We show the utility of our dataset by training an alpha-to-rgb generation method that outperforms baselines. Please see our project website at https://ryanndagreat.github.io/MAGICK/.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationImage MattingText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Deep Video Matting via Spatio-Temporal Alignment and Aggregation

2021-04-22 · CVPR 2021 1 · Yanan sun, Guanzhi Wang, Qiao Gu, Chi-Keung Tang 외

Despite the significant progress made by deep learning in natural image matting, there has been so far no representative work on deep learning for video matting due to the inherent technical challenges in reasoning tempo…

DecoderDeep LearningImage MattingOptical Flow Estimation+1

VideoMaMa: Mask-Guided Video Matting via Generative Prior

2026-01-20 · Sangbeom Lim, Seoung Wug Oh, Jiahui Huang, Heeji Yoon 외 arxiv

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model (VideoMaMa) that converts coarse segmentat…

Zero-shot Generalization

MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator

2025-12-12 · Peiqing Yang, Shangchen Zhou, Kai Hao, Qingyi Tao arxiv

Video matting remains limited by the scale and realism of existing datasets. While leveraging segmentation data can enhance semantic stability, the lack of effective boundary supervision often leads to segmentation-like …

Image Matting

Text-to-Image Synthesis Based on Machine Generated Captions

2019-10-09 · Marco Menardi, Alex Falcon, Saida S. Mohamed, Lorenzo Seidenari 외

Text to Image Synthesis refers to the process of automatic generation of a photo-realistic image starting from a given text and is revolutionizing many real-world applications. In order to perform such process it is nece…

Image CaptioningImage Generation

Latent Multimodal Reconstruction for Misinformation Detection

2025-04-08 · Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis

Multimodal misinformation, such as miscaptioned images, where captions misrepresent an image's origin, context, or meaning, poses a growing challenge in the digital age. To support fact-checkers, researchers have been fo…

Image ReconstructionMisinformation