paper-with-me

홈 › Papers

From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors via LLM-guided Symbolic Reasoning

2025-02-09 · Yuhui Zeng, Haoxiang Wu, Wenjie Nie, Xiawu Zheng, Guangyao Chen, Yunhang Shen, Jun Peng, Yonghong Tian, Rongrong Ji

Our key innovation lies in bridging the semantic gap between object detection and event understanding without requiring expensive task-specific training. The proposed plug-and-play framework interfaces with any open-vocabulary detector while extending their inherent capabilities across architectures. At its core, our approach combines (i) a symbolic regression mechanism exploring relationship patterns among detected entities and (ii) a LLM-guided strategically guiding the search toward meaningful expressions. These discovered symbolic rules transform low-level visual perception into interpretable event understanding, providing a transparent reasoning path from objects to events with strong transferability across domains.We compared our training-free framework against specialized event recognition systems across diverse application domains. Experiments demonstrate that our framework enhances multiple object detector architectures to recognize complex events such as illegal fishing activities (75% AUROC, +8.36% improvement), construction safety violations (+15.77%), and abnormal crowd behaviors (+23.16%). The code will be released soon.

📄 PDF Abstract BibTeX arXiv:2502.05843

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionSymbolic Regression

Similar Papers 제목 키워드 기반

PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models

2026-05-08 · Yuliang Li, Chu Zhou, Heng Guo, Boxin Shi 외 arxiv

Mainstream vision-language models (VLMs) fundamentally struggle with severe optical ambiguities, such as reflections and transparent objects, due to the inherent limitations of standard RGB inputs. While polarization ima…

Debiasing Event Understanding for Visual Commonsense Tasks

2022-05-01 · Findings (ACL) 2022 5 · Minji Seo, YeonJoon Jung, Seungtaek Choi, Seung-won Hwang 외

We study event understanding as a critical step towards visual commonsense tasks.Meanwhile, we argue that current object-based event understanding is purely likelihood-based, leading to incorrect event prediction, due to…

Multimodal Reference Visual Grounding

2025-04-02 · Yangxiao Lu, Ruosen Li, Liqiang Jing, Jikai Wang 외

Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding performance by training large models with …

Few-Shot Object DetectionVisual Grounding

CLEVRER: CoLlision Events for Video REpresentation and Reasoning

2019-10-03 · ICLR 2020 1 · Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 외

The ability to reason about temporal and causal events from videos lies at the core of human intelligence. Most video reasoning benchmarks, however, focus on pattern recognition from complex visual and language input, in…

counterfactualDescriptiveDiagnosticVisual Reasoning

Moments in Time Dataset: one million videos for event understanding

2018-01-09 · Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan 외

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics ev…

Action RecognitionDiversityMultimodal Activity RecognitionTemporal Action Localization