paper-with-me

홈 › Papers

Semantically Invariant Text-to-Image Generation

2018-09-27 · Shagan Sah, Dheeraj Peri, Ameya Shringi, Chi Zhang, Miguel Dominguez, Andreas Savakis, Ray Ptucha

Image captioning has demonstrated models that are capable of generating plausible text given input images or videos. Further, recent work in image generation has shown significant improvements in image quality when text is used as a prior. Our work ties these concepts together by creating an architecture that can enable bidirectional generation of images and text. We call this network Multi-Modal Vector Representation (MMVR). Along with MMVR, we propose two improvements to the text conditioned image generation. Firstly, a n-gram metric based cost function is introduced that generalizes the caption with respect to the image. Secondly, multiple semantically similar sentences are shown to help in generating better images. Qualitative and quantitative evaluations demonstrate that MMVR improves upon existing text conditioned image generation results by over 20%, while integrating visual and text modalities.

📄 PDF Abstract BibTeX arXiv:1809.10274

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Language-agnostic Semantic Consistent Text-to-Image Generation

2022-05-01 · MML (ACL) 2022 5 · SeongJun Jung, Woo Suk Choi, SeongHo Choi, Byoung-Tak Zhang

Recent GAN-based text-to-image generation models have advanced that they can generate photo-realistic images matching semantically with descriptions. However, research on multi-lingual text-to-image generation has not be…

Generative Adversarial NetworkImage GenerationMulti-lingual Text-to-Image GenerationMultilingual Text-to-Image Generation+2

I4VGen: Image as Free Stepping Stone for Text-to-Video Generation

2024-06-04 · Xiefan Guo, Jinlin Liu, Miaomiao Cui, Liefeng Bo 외

Text-to-video generation has trailed behind text-to-image generation in terms of quality and diversity, primarily due to the inherent complexities of spatio-temporal modeling and the limited availability of video-text da…

DiversityImage GenerationText to Image GenerationText-to-Image Generation+2

Self-Supervised Learning of Pretext-Invariant Representations

2019-12-04 · CVPR 2020 6 · Ishan Misra, Laurens van der Maaten

The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations for a large training set of images. Many …

Contrastive Learningobject-detectionObject DetectionRepresentation Learning+3

RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration

2025-09-15 · Zilong Zhang, Chujie Qin, Chunle Guo, Yong Zhang 외 arxiv

This work presents Robust Representation Learning via Adaptive Mask (RAM++), a two-stage framework for all-in-one image restoration. RAM++ integrates high-level semantic understanding with low-level texture generation to…

Representation LearningImage Restoration

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning

2026-06-29 · Xin Zou, Haolin Deng, Yibo Yan, Shuliang Liu 외 arxiv

Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cross-view consistency, \textit{i.e.}, semantically invariant views of…

Reinforcement LearningMultimodal ReasoningData Augmentation