paper-with-me

Papers

KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models

2023-05-28 · Zhiwei Jia, Pradyumna Narayana, Arjun R. Akula, Garima Pruthi, Hao Su, Sugato Basu, Varun Jampani

Image ad understanding is a crucial task with wide real-world applications. Although highly challenging with the involvement of diverse atypical scenes, real-world entities, and reasoning over scene-texts, how to interpret image ads is relatively under-explored, especially in the era of foundational vision-language models (VLMs) featuring impressive generalizability and adaptability. In this paper, we perform the first empirical study of image ad understanding through the lens of pre-trained VLMs. We benchmark and reveal practical challenges in adapting these VLMs to image ad understanding. We propose a simple feature adaptation strategy to effectively fuse multimodal information for image ads and further empower it with knowledge of real-world entities. We hope our study draws more attention to image ad understanding which is broadly relevant to the advertising industry.

📄 PDF Abstract BibTeX arXiv:2305.18373

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer

2025-10-07 · Muhammad Dehan Al Kautsar, Fajri Koto arxiv

Tokenization defines the foundation of multilingual language models by determining how words are represented and shared across languages. However, existing methods often fail to support effective cross-lingual transfer b…

Representation LearningEmotion ClassificationCross-Lingual TransferHate Speech Detection

Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge

2024-07-05 · Yuanze Lin, Yunsheng Li, Dongdong Chen, Weijian Xu 외

In recent years, multimodal large language models (MLLMs) have made significant strides by training on vast high-quality image-text datasets, enabling them to generally understand images well. However, the inherent diffi…

Instance SegmentationOptical Character Recognition (OCR)RAGRetrieval-augmented Generation+3

Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data

2024-07-18 · Wufei Ma, Kai Li, Zhongshi Jiang, Moustafa Meshry 외

Recent video-text foundation models have demonstrated strong performance on a wide variety of downstream video understanding tasks. Can these video-text models genuinely understand the contents of natural videos? Standar…

Language ModellingLarge Language ModelRetrievalVideo Understanding

Data Augmentation Revisited: Rethinking the Distribution Gap between Clean and Augmented Data

2019-09-19 · Zhuoxun He, Lingxi Xie, Xin Chen, Ya zhang 외

Data augmentation has been widely applied as an effective methodology to improve generalization in particular when training deep neural networks. Recently, researchers proposed a few intensive data augmentation technique…

Data Augmentationimage-classificationImage Classificationobject-detection+1

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

2025-11-17 · Jiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen 외 arxiv

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitati…

Reinforcement Learning