paper-with-me

홈 › Papers

VICTR: Visual Information Captured Text Representation for Text-to-Vision Multimodal Tasks

2020-12-01 · COLING 2020 8 · Caren Han, Siqu Long, Siwen Luo, Kunze Wang, Josiah Poon

Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited information in order to fully describe visually realistic images. We propose a new visual contextual text representation for text-to-image multimodal tasks, VICTR, which captures rich visual semantic information of objects from the text input. First, we use the text description as initial input and conduct dependency parsing to extract the syntactic structure and analyse the semantic aspect, including object quantities, to extract the scene graph. Then, we train the extracted objects, attributes, and relations in the scene graph and the corresponding geometric relation information using Graph Convolutional Networks, and it generates text representation which integrates textual and visual semantic information. The text representation is aggregated with word-level and sentence-level embedding to generate both visual contextual word and sentence representation. For the evaluation, we attached VICTR to the state-of-the-art models in text-to-image generation.VICTR is easily added to existing models and improves across both quantitative and qualitative aspects.

📄 PDF Abstract BibTeX

Code (1)

usydnlp/VICTR 공식 구현 pytorch

Tasks

Dependency ParsingSentence

Methods 이 논문이 사용한 방법론

Graph Convolutional Networks 설명 없음

Similar Papers 제목 키워드 기반

VICTR: Visual Information Captured Text Representation for Text-to-Image Multimodal Tasks

2020-10-07 · Soyeon Caren Han, Siqu Long, Siwen Luo, Kunze Wang 외

Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited information in order to fully describe visuall…

Dependency ParsingSentenceText-to-Image Generation

VicTR: Video-conditioned Text Representations for Activity Recognition

2023-04-05 · CVPR 2024 1 · Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo

Vision-Language models (VLMs) have excelled in the image-domain -- especially in zero-shot settings -- thanks to the availability of vast pretraining data (i.e., paired image-text samples). However for videos, such paire…

Action ClassificationActivity RecognitionFormZero-Shot Action Recognition

ViCTr: Vital Consistency Transfer for Pathology Aware Image Synthesis

2025-05-08 · Onkar Susladkar, Gayatri Deshmukh, Yalcin Tur, Ulas Bagci

Synthesizing medical images remains challenging due to limited annotated pathological data, modality domain gaps, and the complexity of representing diffuse pathologies such as liver cirrhosis. Existing methods often str…

8kData AugmentationImage Generation

Scalable Visual Pretraining for Language Intelligence

2026-07-10 · Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao 외 arxiv

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset…

Learning Spatial-context-aware Global Visual Feature Representation for Instance Image Retrieval

2023-01-01 · ICCV 2023 1 · Zhongyan Zhang, Lei Wang, Luping Zhou, Piotr Koniusz

In instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Neverth…

Image RetrievalRetrieval