Vision and Language Integration: Moving beyond Objects
Code (0)
등록된 구현이 없습니다.
Tasks
Action ClassificationImage CaptioningQuestion AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Unsupervised object-centric video generation and decomposition in 3D
A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying…
3D Object DetectionDepth EstimationDepth PredictionInstance Segmentation+4Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation
We address the unsupervised learning of several interconnected problems in low-level vision: single view depth prediction, camera motion estimation, optical flow, and segmentation of a video into the static scene and mov…
Depth EstimationDepth PredictionMonocular Depth EstimationMotion Estimation+2Beyond Categories: The Visual Memex Model for Reasoning About Object Relationships
The use of context is critical for scene understanding in computer vision, where the recognition of an object is driven by both local appearance and the objects relationship to other elements of the scene (context). Mos…
ObjectScene UnderstandingSpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geomet…
Spatial ReasoningU2-ONet: A Two-level Nested Octave U-structure with Multiscale Attention Mechanism for Moving Instances Segmentation
Most scenes in practical applications are dynamic scenes containing moving objects, so segmenting accurately moving objects is crucial for many computer vision applications. In order to efficiently segment out all moving…