paper-with-me

홈 › Papers

Vision and Language Integration: Moving beyond Objects

2017-01-01 · WS 2017 1 · Ravi Shekhar, S Pezzelle, ro, Aur{\'e}lie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationImage CaptioningQuestion AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Unsupervised object-centric video generation and decomposition in 3D

2020-07-07 · NeurIPS 2020 12 · Paul Henderson, Christoph H. Lampert

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying…

3D Object DetectionDepth EstimationDepth PredictionInstance Segmentation+4

Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation

2018-05-24 · CVPR 2019 6 · Anurag Ranjan, Varun Jampani, Lukas Balles, Kihwan Kim 외

We address the unsupervised learning of several interconnected problems in low-level vision: single view depth prediction, camera motion estimation, optical flow, and segmentation of a video into the static scene and mov…

Depth EstimationDepth PredictionMonocular Depth EstimationMotion Estimation+2

Beyond Categories: The Visual Memex Model for Reasoning About Object Relationships

2009-12-01 · NeurIPS 2009 12 · Tomasz Malisiewicz, Alyosha Efros

The use of context is critical for scene understanding in computer vision, where the recognition of an object is driven by both local appearance and the objects relationship to other elements of the scene (context). Mos…

ObjectScene Understanding

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

2026-03-28 · Jian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi 외 arxiv

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geomet…

Spatial Reasoning

U2-ONet: A Two-level Nested Octave U-structure with Multiscale Attention Mechanism for Moving Instances Segmentation

2020-07-26 · Chenjie Wang, Chengyuan Li, Bin Luo

Most scenes in practical applications are dynamic scenes containing moving objects, so segmenting accurately moving objects is crucial for many computer vision applications. In order to efficiently segment out all moving…