paper-with-me

홈 › Papers

Conditional Image-Text Embedding Networks

2017-11-22 · ECCV 2018 9 · Bryan A. Plummer, Paige Kordas, M. Hadi Kiapour, Shuai Zheng, Robinson Piramuthu, Svetlana Lazebnik

This paper presents an approach for grounding phrases in images which jointly learns multiple text-conditioned embeddings in a single end-to-end model. In order to differentiate text phrases into semantically distinct subspaces, we propose a concept weight branch that automatically assigns phrases to embeddings, whereas prior works predefine such assignments. Our proposed solution simplifies the representation requirements for individual embeddings and allows the underrepresented concepts to take advantage of the shared representations before feeding them into concept-specific layers. Comprehensive experiments verify the effectiveness of our approach across three phrase grounding datasets, Flickr30K Entities, ReferIt Game, and Visual Genome, where we obtain a (resp.) 4%, 3%, and 4% improvement in grounding performance over a strong region-phrase embedding baseline.

📄 PDF Abstract BibTeX arXiv:1711.08389

Code (1)

BryanPlummer/cite 공식 구현 tf

Tasks

Phrase Grounding

Similar Papers 제목 키워드 기반

Visual Chain-of-Thought Diffusion Models

2023-03-28 · William Harvey, Frank Wood

Recent progress with conditional image diffusion models has been stunning, and this holds true whether we are speaking about models conditioned on a text description, a scene layout, or a sketch. Unconditional image diff…

Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models

2025-12-26 · Masayuki Kawarada, Kosuke Yamada, Antonio Tejero-de-Pablos, Naoto Inoue arxiv

Conditional image embeddings are feature representations that focus on specific aspects of an image indicated by a given textual condition (e.g., color, genre), which has been a challenging problem. Although recent visio…

Watch What You Just Said: Image Captioning with Text-Conditional Attention

2016-06-15 · Luowei Zhou, Chenliang Xu, Parker Koch, Jason J. Corso

Attention mechanisms have attracted considerable interest in image captioning due to its powerful performance. However, existing methods use only visual content as attention and whether textual context can improve attent…

Image CaptioningLanguage ModelingLanguage Modelling

Context-Aware Autoregressive Models for Multi-Conditional Image Generation

2025-05-18 · Yixiao Chen, Zhiyuan Ma, Guoli Jia, Che Jiang 외

Autoregressive transformers have recently shown impressive image generation quality and efficiency on par with state-of-the-art diffusion models. Unlike diffusion architectures, autoregressive models can naturally incorp…

Conditional Image GenerationImage Generation

Text-conditional Attribute Alignment across Latent Spaces for 3D Controllable Face Image Synthesis

2024-01-01 · CVPR 2024 1 · Feifan Xu, Rui Li, Si Wu, Yong Xu 외

With the advent of generative models and vision language pretraining significant improvement has been made in text-driven face manipulation. The text embedding can be used as target supervision for expression control…

AttributeImage Generation