paper-with-me

Papers

TokenMotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection Via Learnable Token Selection

2023-11-05 · Zifan Yu, Erfan Bank Tavakoli, Meida Chen, Suya You, Raghuveer Rao, Sanjeev Agarwal, Fengbo Ren

The area of Video Camouflaged Object Detection (VCOD) presents unique challenges in the field of computer vision due to texture similarities between target objects and their surroundings, as well as irregular motion patterns caused by both objects and camera movement. In this paper, we introduce TokenMotion (TMNet), which employs a transformer-based model to enhance VCOD by extracting motion-guided features using a learnable token selection. Evaluated on the challenging MoCA-Mask dataset, TMNet achieves state-of-the-art performance in VCOD. It outperforms the existing state-of-the-art method by a 12.8% improvement in weighted F-measure, an 8.4% enhancement in S-measure, and a 10.7% boost in mean IoU. The results demonstrate the benefits of utilizing motion-guided features via learnable token selection within a transformer-based framework to tackle the intricate task of VCOD.

📄 PDF Abstract BibTeX arXiv:2311.02535

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject Detection

Similar Papers 제목 키워드 기반

TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation

2025-04-11 · CVPR 2025 1 · Ruineng Li, Daitao Xing, Huiming Sun, Yuanzhou Ha 외

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video…

DisentanglementVideo Generation

Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces

2025-12-11 · Bishoy Galoaa, Xiangyu Bai, Sarah Ostadabbas arxiv

We present Lang2Motion, a framework for language-guided point trajectory generation by aligning motion manifolds with joint embedding spaces. Unlike prior work focusing on human motion or video synthesis, we generate exp…

Action RecognitionVideo GenerationPoint TrackingStyle Transfer

2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation

2024-06-20 · Bin Cao, Yisi Zhang, Xuanxu Lin, Xingjian He 외

Motion Expression guided Video Segmentation is a challenging task that aims at segmenting objects in the video based on natural language expressions with motion descriptions. Unlike the previous referring video object se…

Instance SegmentationReferring Video Object SegmentationSegmentationSemantic Segmentation+4

Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer

2024-03-20 · Yu Deng, Duomin Wang, Baoyuan Wang

In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseu…

MiVE: Multiscale Vision-language features for reference-guided video Editing

2026-05-14 · Tong Wang, Meng Zou, Chengjing Wu, Xiaochao Qu 외 arxiv

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content…