paper-with-me

홈 › Papers

Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding

2025-08-28 · Gowreesh Mago, Pascal Mettes, Stevan Rudinac arxiv

The automatic understanding of video content is advancing rapidly. Empowered by deeper neural networks and large datasets, machines are increasingly capable of understanding what is concretely visible in video frames, whether it be objects, actions, events, or scenes. In comparison, humans retain a unique ability to also look beyond concrete entities and recognize abstract concepts like justice, freedom, and togetherness. Abstract concept recognition forms a crucial open challenge in video understanding, where reasoning on multiple semantic levels based on contextual information is key. In this paper, we argue that the recent advances in foundation models make for an ideal setting to address abstract concept understanding in videos. Automated understanding of high-level abstract concepts is imperative as it enables models to be more aligned with human reasoning and values. In this survey, we study different tasks and datasets used to understand abstract concepts in video content. We observe that, periodically and over a long period, researchers have attempted to solve these tasks, making the best use of the tools available at their disposal. We advocate that drawing on decades of community experience will help us shed light on this important open grand challenge and avoid ``re-inventing the wheel'' as we start revisiting it in the era of multi-modal foundation models.

📄 PDF Abstract BibTeX arXiv:2508.20765

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

2026-05-14 · Mor Ventura, Roy Hirsch, Yonatan Bitton, Regev Cohen 외 arxiv

Humans naturally communicate through abstract concepts like "mood". However, current image editing benchmarks focus primarily on explicit, literal commands, leaving abstract instructions largely underexplored. In this wo…

Image Editing

A Survey of Current Datasets for Vision and Language Research

2015-06-23 · EMNLP 2015 9 · Francis Ferraro, Nasrin Mostafazadeh, Ting-Hao, Huang 외

Integrating vision and language has long been a dream in work on artificial intelligence (AI). In the past two years, we have witnessed an explosion of work that brings together vision and language from images to videos …

Survey

Interpreting Language Models Through Concept Descriptions: A Survey

2025-10-01 · Nils Feldhus, Laura Kopf arxiv

Understanding the decision-making processes of neural networks is a central goal of mechanistic interpretability. In the context of Large Language Models (LLMs), this involves uncovering the underlying mechanisms and ide…

Ambiguity invokes Creativity : looking through Quantum physics

2017-02-23

Creativity, defined as the tendency to generate or recognize new ideas or alternatives and to make connections between seemingly unrelated phenomena, is too vast a horizon to be summed up in such a simple sentence. The e…

Sentence

Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories

2023-08-21 · Delfina Sol Martinez Pandiani, Valentina Presutti

The field of Computer Vision (CV) is increasingly shifting towards ``high-level'' visual sensemaking tasks, yet the exact nature of these tasks remains unclear and tacit. This survey paper addresses this ambiguity by sys…

ClassificationClusteringimage-classificationImage Classification+2