VICTR: Visual Information Captured Text Representation for Text-to-Vision Multimodal Tasks
Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited information in order to fully describe visually realistic images. We propose a new visual contextual text representation for text-to-image multimodal tasks, VICTR, which captures rich visual semantic information of objects from the text input. First, we use the text description as initial input and conduct dependency parsing to extract the syntactic structure and analyse the semantic aspect, including object quantities, to extract the scene graph. Then, we train the extracted objects, attributes, and relations in the scene graph and the corresponding geometric relation information using Graph Convolutional Networks, and it generates text representation which integrates textual and visual semantic information. The text representation is aggregated with word-level and sentence-level embedding to generate both visual contextual word and sentence representation. For the evaluation, we attached VICTR to the state-of-the-art models in text-to-image generation.VICTR is easily added to existing models and improves across both quantitative and qualitative aspects.
Code (1)
Tasks
Dependency ParsingSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
VICTR: Visual Information Captured Text Representation for Text-to-Image Multimodal Tasks
Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited information in order to fully describe visuall…
Dependency ParsingSentenceText-to-Image GenerationVicTR: Video-conditioned Text Representations for Activity Recognition
Vision-Language models (VLMs) have excelled in the image-domain -- especially in zero-shot settings -- thanks to the availability of vast pretraining data (i.e., paired image-text samples). However for videos, such paire…
Action ClassificationActivity RecognitionFormZero-Shot Action RecognitionViCTr: Vital Consistency Transfer for Pathology Aware Image Synthesis
Synthesizing medical images remains challenging due to limited annotated pathological data, modality domain gaps, and the complexity of representing diffuse pathologies such as liver cirrhosis. Existing methods often str…
8kData AugmentationImage GenerationScalable Visual Pretraining for Language Intelligence
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset…
Learning Spatial-context-aware Global Visual Feature Representation for Instance Image Retrieval
In instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Neverth…
Image RetrievalRetrieval