TextManiA: Enriching Visual Feature by Text-driven Manifold Augmentation
We propose TextManiA, a text-driven manifold augmentation method that semantically enriches visual feature spaces, regardless of class distribution. TextManiA augments visual data with intra-class semantic perturbation by exploiting easy-to-understand visually mimetic words, i.e., attributes. This work is built on an interesting hypothesis that general language models, e.g., BERT and GPT, encompass visual information to some extent, even without training on visual training data. Given the hypothesis, TextManiA transfers pre-trained text representation obtained from a well-established large language encoder to a target visual feature space being learned. Our extensive analysis hints that the language encoder indeed encompasses visual information at least useful to augment visual representation. Our experiments demonstrate that TextManiA is particularly powerful in scarce samples with class imbalance as well as even distribution. We also show compatibility with the label mix-based approaches in evenly distributed scarce data.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Semantic-enriched Visual Vocabulary Construction in a Weakly Supervised Context
One of the prevalent learning tasks involving images is content-based image classification. This is a difficult task especially because the low-level features used to digitally describe images usually capture little info…
ClassificationGeneral Classificationimage-classificationImage ClassificationAction2Dialogue: Generating Character-Centric Narratives from Scene-Level Prompts
Recent advances in scene-based video generation have enabled systems to synthesize coherent visual narratives from structured prompts. However, a crucial dimension of storytelling -- character-driven dialogue and speech …
Dialogue GenerationLarge Language ModelStory GenerationVideo Generation+1Enriching Local and Global Contexts for Temporal Action Localization
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and suff…
Action ClassificationAction LocalizationRetrievalTemporal Action Localization+1Enriching Texture Analysis with Semantic Data
We argue for the importance of explicit semantic modelling in human-centred texture analysis tasks such as retrieval, annotation, synthesis, and zero-shot learning. To this end, low-level attributes are selected and used…
feature selectionRetrievalTexture ClassificationZero-Shot LearningFaceInsight: A Multimodal Large Language Model for Face Perception
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, ofte…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model