Dynamic Scene Understanding from Vision-Language Representations
Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their interactions. Current approaches use distinct methods tailored to sub-tasks such as Situation Recognition and detection of Human-Human and Human-Object Interactions. However, recent advances in image understanding have often leveraged web-scale vision-language (V&L) representations to obviate task-specific engineering. In this work, we propose a framework for dynamic scene understanding tasks by leveraging knowledge from modern, frozen V&L representations. By framing these tasks in a generic manner - as predicting and parsing structured text, or by directly concatenating representations to the input of existing models - we achieve state-of-the-art results while using a minimal number of trainable parameters relative to existing approaches. Moreover, our analysis of dynamic knowledge of these representations shows that recent, more powerful representations effectively encode dynamic scene semantics, making this approach newly possible.
Code (0)
등록된 구현이 없습니다.
Tasks
Grounded Situation RecognitionHuman-Human Interaction RecognitionHuman Interaction RecognitionHuman-Object Interaction DetectionScene UnderstandingSituation RecognitionSimilar Papers 제목 키워드 기반
CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds w…
Contrastive LearningPoint CloudsOmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
Human vision is capable of transforming two-dimensional observations into an egocentric three-dimensional scene understanding, which underpins the ability to translate complex scenes and exhibit adaptive behaviors. This …
Visual Question AnsweringKnowledge DistillationScene UnderstandingAutonomous DrivingObject-Centric Representation Learning with Generative Spatial-Temporal Factorization
Learning object-centric scene representations is essential for attaining structural understanding and abstraction of complex scenes. Yet, as current approaches for unsupervised object-centric representation learning are …
ObjectRepresentation LearningLLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
Developing a multi-modal language model capable of understanding 3D scenes remains challenging due to the limited availability of 3D training data, in contrast to the abundance of 2D datasets used for vision-language mod…
Multi-View 3D ReconstructionScene UnderstandingSLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes
We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling…
Dynamic ReconstructionScene Parsing