paper-with-me

Papers

Fine-grained Visual-textual Representation Learning

2017-08-31 · Xiangteng He, Yuxin Peng

Fine-grained visual categorization is to recognize hundreds of subcategories belonging to the same basic-level category, which is a highly challenging task due to the quite subtle and local visual distinctions among similar subcategories. Most existing methods generally learn part detectors to discover discriminative regions for better categorization performance. However, not all parts are beneficial and indispensable for visual categorization, and the setting of part detector number heavily relies on prior knowledge as well as experimental validation. As is known to all, when we describe the object of an image via textual descriptions, we mainly focus on the pivotal characteristics, and rarely pay attention to common characteristics as well as the background areas. This is an involuntary transfer from human visual attention to textual attention, which leads to the fact that textual attention tells us how many and which parts are discriminative and significant to categorization. So textual attention could help us to discover visual attention in image. Inspired by this, we propose a fine-grained visual-textual representation learning (VTRL) approach, and its main contributions are: (1) Fine-grained visual-textual pattern mining devotes to discovering discriminative visual-textual pairwise information for boosting categorization performance through jointly modeling vision and text with generative adversarial networks (GANs), which automatically and adaptively discovers discriminative parts. (2) Visual-textual representation learning jointly combines visual and textual information, which preserves the intra-modality and inter-modality information to generate complementary fine-grained representation, as well as further improves categorization performance.

📄 PDF Abstract BibTeX arXiv:1709.00340

Code (1)

PKU-ICST-MIPL/OPAM_TIP2018

Tasks

Fine-Grained Visual CategorizationRepresentation Learning

Similar Papers 제목 키워드 기반

Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval

2025-04-10 · Zehong Ma, Hao Chen, Wei Zeng, Limin Su 외

Fine-grained text-to-image retrieval aims to retrieve a fine-grained target image with a given text query. Existing methods typically assume that each training image is accurately depicted by its textual descriptions. Ho…

Image RetrievalRepresentation LearningRetrieval

Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual Features

2020-01-14 · Andres Mafla, Sounak Dey, Ali Furkan Biten, Lluis Gomez 외

Text contained in an image carries high-level semantics that can be exploited to achieve richer image understanding. In particular, the mere presence of text provides strong guiding content that should be employed to tac…

ClassificationDiversityFine-Grained Image ClassificationGeneral Classification+7

FILIP: Fine-grained Interactive Language-Image Pre-Training

2021-11-09 · ICLR 2022 4 · Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu 외

Unsupervised large-scale vision-language pre-training has shown promising advances on various downstream tasks. Existing methods often model the cross-modal interaction either via the similarity of the global feature of …

image-classificationImage ClassificationImage-text RetrievalRetrieval+2

VGSG: Vision-Guided Semantic-Group Network for Text-based Person Search

2023-11-13 · Shuting He, Hao Luo, Wei Jiang, Xudong Jiang 외

Text-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing m…

Person SearchText based Person RetrievalText based Person SearchTransfer Learning

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning

2026-07-02 · Xuehui Wang, Xuankun Yang, Wei Shen arxiv

Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing methods often fail to preserve critical cues under dense instructions and fine-grained queries. In thi…