paper-with-me

홈 › Papers

TextManiA: Enriching Visual Feature by Text-driven Manifold Augmentation

2023-07-27 · ICCV 2023 1 · Moon Ye-Bin, Jisoo Kim, Hongyeob Kim, Kilho Son, Tae-Hyun Oh

We propose TextManiA, a text-driven manifold augmentation method that semantically enriches visual feature spaces, regardless of class distribution. TextManiA augments visual data with intra-class semantic perturbation by exploiting easy-to-understand visually mimetic words, i.e., attributes. This work is built on an interesting hypothesis that general language models, e.g., BERT and GPT, encompass visual information to some extent, even without training on visual training data. Given the hypothesis, TextManiA transfers pre-trained text representation obtained from a well-established large language encoder to a target visual feature space being learned. Our extensive analysis hints that the language encoder indeed encompasses visual information at least useful to augment visual representation. Our experiments demonstrate that TextManiA is particularly powerful in scarce samples with class imbalance as well as even distribution. We also show compatibility with the label mix-based approaches in evenly distributed scarce data.

📄 PDF Abstract BibTeX arXiv:2307.14611

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
WordPiece 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Semantic-enriched Visual Vocabulary Construction in a Weakly Supervised Context

2015-12-14 · Marian-Andrei Rizoiu, Julien Velcin, Stéphane Lallich

One of the prevalent learning tasks involving images is content-based image classification. This is a difficult task especially because the low-level features used to digitally describe images usually capture little info…

ClassificationGeneral Classificationimage-classificationImage Classification

Action2Dialogue: Generating Character-Centric Narratives from Scene-Level Prompts

2025-05-22 · Taewon Kang, Ming C. Lin

Recent advances in scene-based video generation have enabled systems to synthesize coherent visual narratives from structured prompts. However, a crucial dimension of storytelling -- character-driven dialogue and speech …

Dialogue GenerationLarge Language ModelStory GenerationVideo Generation+1

Enriching Local and Global Contexts for Temporal Action Localization

2021-07-27 · ICCV 2021 10 · Zixin Zhu, Wei Tang, Le Wang, Nanning Zheng 외

Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and suff…

Action ClassificationAction LocalizationRetrievalTemporal Action Localization+1

Enriching Texture Analysis with Semantic Data

2013-06-01 · CVPR 2013 6 · Tim Matthews, Mark S. Nixon, Mahesan Niranjan

We argue for the importance of explicit semantic modelling in human-centred texture analysis tasks such as retrieval, annotation, synthesis, and zero-shot learning. To this end, low-level attributes are selected and used…

feature selectionRetrievalTexture ClassificationZero-Shot Learning

FaceInsight: A Multimodal Large Language Model for Face Perception

2025-04-22 · Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang 외

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, ofte…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model