Papers Video Object Tracking
“Video Object Tracking” 태그가 달린 논문 107편 · 필터 해제
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense…
Spatio-Temporal Video GroundingVideo Object TrackingRethinking Object-Centric Representations for Video Dynamics Modeling
Unsupervised video object tracking aims to decompose dynamic scenes into persistent, object-centric entities without manual annotations. Many recent approaches rely on slot-based representations, where a fixed set of lat…
Video Object TrackingTetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking
Track materialization converts raw videos into reusable object tracks that downstream queries can run against without rerunning tracking, but extracting those tracks efficiently and with high fidelity remains expensive. …
Video Object TrackingSiamGM: Siamese Geometry-Aware and Motion-Guided Network for Real-Time Satellite Video Object Tracking
Single object tracking in satellite videos is inherently challenged by small target, blurred background, large aspect ratio changes, and frequent visual occlusions. These constraints often cause appearance-based trackers…
Video Object TrackingV$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across different viewpoints (e.g., egocentric and ex…
Video Object TrackingSatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
Existing satellite video tracking methods often struggle with generalization, requiring scenario-specific training to achieve satisfactory performance, and are prone to track loss in the presence of occlusion. To address…
Video Object TrackingMOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
Video object segmentation (VOS) aims to segment specified target objects throughout a video. Although state-of-the-art methods have achieved impressive performance (e.g., 90+% J&F) on benchmarks such as DAVIS and YouTube…
Video Object SegmentationVideo Object TrackingDepthwise-Dilated Convolutional Adapters for Medical Object Tracking and Segmentation Using the Segment Anything Model 2
Recent advances in medical image segmentation have been driven by deep learning; however, most existing methods remain limited by modality-specific designs and exhibit poor adaptability to dynamic medical imaging scenari…
Medical Image SegmentationVideo Object TrackingVideo SegmentationTumor SegmentationOpenHuman4D: Open-Vocabulary 4D Human Parsing
Understanding dynamic 3D human representation has become increasingly critical in virtual and extended reality applications. However, existing human part segmentation methods are constrained by reliance on closed-set dat…
Human Part SegmentationVideo Object TrackingHuman ParsingHiM2SAM: Enhancing SAM2 with Hierarchical Motion Estimation and Memory Optimization towards Long-term Tracking
This paper presents enhancements to the SAM2 framework for video object tracking task, addressing challenges such as occlusions, background clutter, and target reappearance. We introduce a hierarchical motion estimation …
Motion EstimationObject TrackingVideo Object TrackingEnhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
Successful video analysis relies on accurate recognition of pixels across frames, and frame reconstruction methods based on video correspondence learning are popular due to their efficiency. Existing frame reconstruction…
Decision MakingObjectObject TrackingSemantic Segmentation+1Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusi…
MambaObject TrackingRgb-T TrackingVideo Object TrackingExploring Enhanced Contextual Information for Video-Level Object Tracking
Contextual information at the video level has become increasingly crucial for visual object tracking. However, existing methods typically use only a few tokens to convey this information, which can lead to information lo…
ObjectObject TrackingSemi-Supervised Video Object SegmentationVideo Object Tracking+2Referring Video Object Segmentation via Language-aligned Track Selection
Referring video object segmentation (RVOS) requires tracking and segmenting an object throughout a video according to a given natural language expression, demanding both complex motion understanding and the alignment of …
ObjectObject TrackingReferring Video Object SegmentationSemantic Segmentation+3Teaching VLMs to Localize Specific Objects from In-context Examples
Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks.…
ObjectObject TrackingQuestion AnsweringVideo Object Tracking+3NT-VOT211: A Large-Scale Benchmark for Night-time Visual Object Tracking
Many current visual object tracking benchmarks such as OTB100, NfS, UAV123, LaSOT, and GOT-10K, predominantly contain day-time scenarios while the challenges posed by the night-time has been less investigated. It is prim…
Object TrackingVideo Object TrackingVisual Object TrackingDepth Attention for Robust RGB Tracking
RGB video object tracking is a fundamental task in computer vision. Its effectiveness can be improved using depth information, particularly for handling motion-blurred target. However, depth information is often missing …
Depth EstimationMonocular Depth EstimationObject TrackingVideo Object TrackingVOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking
Open-vocabulary multi-object tracking (OVMOT) represents a critical new challenge involving the detection and tracking of diverse object categories in videos, encompassing both seen categories (base classes) and unseen c…
Multi-Object TrackingObjectobject-detectionObject Detection+4Associate Everything Detected: Facilitating Tracking-by-Detection to the Unknown
Multi-object tracking (MOT) emerges as a pivotal and highly promising branch in the field of computer vision. Classical closed-vocabulary MOT (CV-MOT) methods aim to track objects of predefined categories. Recently, some…
Multi-Object TrackingMultiple Object TrackingObjectObject Tracking+1Medical SAM 2: Segment medical images as video via Segment Anything Model 2
Medical image segmentation plays a pivotal role in clinical diagnostics and treatment planning, yet existing models often face challenges in generalization and in handling both 2D and 3D data uniformly. In this paper, we…
Image SegmentationInteractive SegmentationMedical Image SegmentationObject Tracking+4