paper-with-me

홈 › Papers

Track Everything Everywhere Fast and Robustly

2024-03-26 · Yunzhou Song, Jiahui Lei, ZiYun Wang, Lingjie Liu, Kostas Daniilidis

We propose a novel test-time optimization approach for efficiently and robustly tracking any pixel at any time in a video. The latest state-of-the-art optimization-based tracking technique, OmniMotion, requires a prohibitively long optimization time, rendering it impractical for downstream applications. OmniMotion is sensitive to the choice of random seeds, leading to unstable convergence. To improve efficiency and robustness, we introduce a novel invertible deformation network, CaDeX++, which factorizes the function representation into a local spatial-temporal feature grid and enhances the expressivity of the coupling blocks with non-linear functions. While CaDeX++ incorporates a stronger geometric bias within its architectural design, it also takes advantage of the inductive bias provided by the vision foundation models. Our system utilizes monocular depth estimation to represent scene geometry and enhances the objective by incorporating DINOv2 long-term semantics to regulate the optimization process. Our experiments demonstrate a substantial improvement in training speed (more than \textbf{10 times} faster), robustness, and accuracy in tracking over the SoTA optimization-based method OmniMotion.

📄 PDF Abstract BibTeX arXiv:2403.17931

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationInductive BiasMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Tracking Everything Everywhere All at Once

2023-06-08 · ICCV 2023 1 · Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li 외

We present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited temporal windows,…

AllMotion EstimationOptical Flow Estimation

Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense

2024-11-22 · Jie Zhang, Christian Schlarmann, Kristina Nikolić, Nicholas Carlini 외

Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations at multiple noisy i…

Adversarial AttackRobust classification

Decomposition Betters Tracking Everything Everywhere

2024-07-09 · Rui Li, Dong Liu

Recent studies on motion estimation have advocated an optimized motion representation that is globally consistent across the entire video, preferably for every pixel. This is challenging as a uniform representation may n…

Motion EstimationPoint Tracking

Segment Everything Everywhere All at Once

2023-04-13 · NeurIPS 2023 11 · Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li 외

In this work, we present SEEM, a promptable and interactive model for segmenting everything everywhere all at once in an image, as shown in Fig.1. In SEEM, we propose a novel decoding mechanism that enables diverse promp…

AllDecoderImage SegmentationInteractive Segmentation+5

Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual Samples

2023-01-01 · ICCV 2023 1 · Mingfei Chen, Kun Su, Eli Shlizerman

Fully immersive and interactive audio-visual scenes are dynamic such that the listeners and the sound emitters move and interact with each other. Reconstruction of an immersive sound experience, as it happens in the …