paper-with-me

홈 › Papers

Identity-Aware Textual-Visual Matching with Latent Co-attention

2017-08-07 · ICCV 2017 10 · Shuang Li, Tong Xiao, Hongsheng Li, Wei Yang, Xiaogang Wang

Textual-visual matching aims at measuring similarities between sentence descriptions and images. Most existing methods tackle this problem without effectively utilizing identity-level annotations. In this paper, we propose an identity-aware two-stage framework for the textual-visual matching problem. Our stage-1 CNN-LSTM network learns to embed cross-modal features with a novel Cross-Modal Cross-Entropy (CMCE) loss. The stage-1 network is able to efficiently screen easy incorrect matchings and also provide initial training point for the stage-2 training. The stage-2 CNN-LSTM network refines the matching results with a latent co-attention mechanism. The spatial attention relates each word with corresponding image regions while the latent semantic attention aligns different sentence structures to make the matching results more robust to sentence structure variations. Extensive experiments on three datasets with identity-level annotations show that our framework outperforms state-of-the-art approaches by large margins.

📄 PDF Abstract BibTeX arXiv:1708.01988

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceText based Person Retrieval

Similar Papers 제목 키워드 기반

FlexID: Training-Free Flexible Identity Injection via Intent-Aware Modulation for Text-to-Image Generation

2026-02-07 · Guandong Li, Yijun Ding arxiv

Personalized text-to-image generation aims to seamlessly integrate specific identities into textual descriptions. However, existing training-free methods often rely on rigid visual feature injection, creating a conflict …

Text-to-Image Generation

Face Time Traveller : Travel Through Ages Without Losing Identity

2026-02-26 · Purbayan Kar, Ayush Ghadiya, Vishal Chudasama, Pankaj Wasnik 외 arxiv

Face aging, an ill-posed problem shaped by environmental and genetic factors, is vital in entertainment, forensics, and digital archiving, where realistic age transformations must preserve both identity and visual realis…

Taming Identity Consistency and Prompt Diversity in Diffusion Models via Latent Concatenation and Masked Conditional Flow Matching

2025-11-11 · Aditi Singhania, Arushi Jain, Krutik Malani, Riddhi Dhawan 외 arxiv

Subject-driven image generation aims to synthesize novel depictions of a specific subject across diverse contexts while preserving its core identity features. Achieving both strong identity consistency and high prompt di…

parameter-efficient fine-tuningImage Generation

Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

2026-06-01 · Liyuan Ma, Xueji Fang, Guo-Jun Qi arxiv

Image customization learns target subjects from reference concept images and generates conditioned images per text prompts, mainly modifying styles or backgrounds. Prevailing methods adopt fine-tuning to pack diverse con…

Visual-textual Dermatoglyphic Animal Biometrics: A First Case Study on Panthera tigris

2025-12-16 · Wenshuo Li, Majid Mirmehdi, Tilo Burghardt arxiv

Biologists have long combined visuals with textual field notes to re-identify (Re-ID) animals. Contemporary AI tools automate this for species with distinctive morphological features but remain largely image-based. Here,…

Cross-Modal Retrieval