Ad Lingua: Text Classification Improves Symbolism Prediction in Image Advertisements
Understanding image advertisements is a challenging task, often requiring non-literal interpretation. We argue that standard image-based predictions are insufficient for symbolism prediction. Following the intuition that texts and images are complementary in advertising, we introduce a multimodal ensemble of a state of the art image-based classifier, a classifier based on an object detection architecture, and a fine-tuned language model applied to texts extracted from ads by OCR. The resulting system establishes a new state of the art in symbolism prediction.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage Modellingobject-detectionObject DetectionOptical Character Recognition (OCR)Predictiontext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Marriage is a Peach and a Chalice: Modelling Cultural Symbolism on the SemanticWeb
In this work, we fill the gap in the Semantic Web in the context of Cultural Symbolism. Building upon earlier work in, we introduce the Simulation Ontology, an ontology that models the background knowledge of symbolic me…
Dyr Bul Shchyl. Proxying Sound Symbolism With Word Embeddings
This paper explores modern word embeddings in the context of sound symbolism. Using basic properties of the representations space one can construct semantic axes. A method is proposed to measure if the presence of indivi…
Word EmbeddingsAdversarially Probing Cross-Family Sound Symbolism in 27 Languages
The phenomenon of sound symbolism, the non-arbitrary mapping between word sounds and meanings, has long been demonstrated through anecdotal experiments like Bouba Kiki, but rarely tested at scale. We present the first co…
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs long natural-language rationales that are poorly aligned with efficient machine reasoning. We propose Communi…
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
Recently, Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated aptitude as potential substitutes for human participants in experiments testing psycholinguistic phenomena. However, an understud…