MMC: Multi-Modal Colorization of Images using Textual Descriptions
Handling various objects with different colors is a significant challenge for image colorization techniques. Thus, for complex real-world scenes, the existing image colorization algorithms often fail to maintain color consistency. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To do so, we have proposed a deep network that takes two inputs (grayscale image and the respective encoded text description) and tries to predict the relevant color components. Also, we have predicted each object in the image and have colorized them with their individual description to incorporate their specific attributes in the colorization process. After that, a fusion model fuses all the image objects (segments) to generate the final colorized image. As the respective textual descriptions contain color information of the objects present in the image, text encoding helps to improve the overall quality of predicted colors. In terms of performance, the proposed method outperforms existing colorization techniques in terms of LPIPS, PSNR and SSIM metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
ColorizationImage ColorizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Language-based Image Colorization: A Benchmark and Beyond
Image colorization aims to bring colors back to grayscale images. Automatic image colorization methods, which requires no additional guidance, struggle to generate high-quality images due to color ambiguity, and provides…
BenchmarkingColorizationcross-modal alignmentImage ColorizationTIC: Text-Guided Image Colorization
Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make…
ColorizationImage ColorizationL-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors
Language-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color des…
ColorizationStructurally Consistent MRI Colorization using Cross-modal Fusion Learning
Medical image colorization can greatly enhance the interpretability of the underlying imaging modality and provide insights into human anatomy. The objective of medical image colorization is to transfer a diverse spectru…
AnatomyColorizationFeature CompressionImage ColorizationLanguage-based Colorization of Scene Sketches
Being natural, touchless, and fun-embracing, language-based inputs have been demonstrated effective for various tasks from image generation to literacy education for children. This paper for the first time presents a lan…
ColorizationImage GenerationScene UnderstandingSketch