paper-with-me

Papers

Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor Segmentation

2021-05-14 · CVPR 2021 1 · Tianrui Hui, Shaofei Huang, Si Liu, Zihan Ding, Guanbin Li, Wenguan Wang, Jizhong Han, Fei Wang

Language-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a general encoder to extract a mixed spatio-temporal feature for the target frame. Though 3D convolutions are amenable to recognizing which actor is performing the queried actions, it also inevitably introduces misaligned spatial information from adjacent frames, which confuses features of the target frame and yields inaccurate segmentation. Therefore, we propose a collaborative spatial-temporal encoder-decoder framework which contains a 3D temporal encoder over the video clip to recognize the queried actions, and a 2D spatial encoder over the target frame to accurately segment the queried actors. In the decoder, a Language-Guided Feature Selection (LGFS) module is proposed to flexibly integrate spatial and temporal features from the two encoders. We also propose a Cross-Modal Adaptive Modulation (CMAM) module to dynamically recombine spatial- and temporal-relevant linguistic features for multimodal feature interaction in each stage of the two encoders. Our method achieves new state-of-the-art performance on two popular benchmarks with less computational overhead than previous approaches.

📄 PDF Abstract BibTeX arXiv:2105.06818

Code (0)

등록된 구현이 없습니다.

Tasks

Decoderfeature selectionReferring Expression Segmentation

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

CTM: Collaborative Temporal Modeling for Action Recognition

2020-02-08 · Qian Liu, Tao Wang, Jie Liu, Yang Guan 외

With the rapid development of digital multimedia, video understanding has become an important field. For action recognition, temporal dimension plays an important role, and this is quite different from image recognition.…

Action RecognitionVideo Understanding

Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding

2023-01-01 · CVPR 2023 1 · Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin 외

Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic vis…

ObjectSpatio-Temporal Video GroundingVideo Grounding

Region-aware Spatiotemporal Modeling with Collaborative Domain Generalization for Cross-Subject EEG Emotion Recognition

2026-01-22 · Weiwei Wu, Yueyang Li, Yuhu Shi, Weiming Zeng 외 arxiv

Cross-subject EEG-based emotion recognition (EER) remains challenging due to strong inter-subject variability, which induces substantial distribution shifts in EEG signals, as well as the high complexity of emotion-relat…

EEG Emotion RecognitionDomain Generalization

CollaMamba: Efficient Collaborative Perception with Cross-Agent Spatial-Temporal State Space Model

2024-09-12 · Yang Li, Quan Yuan, Guiyang Luo, Xiaoyuan Fu 외

By sharing complementary perceptual information, multi-agent collaborative perception fosters a deeper understanding of the environment. Recent studies on collaborative perception mostly utilize CNNs or Transformers to l…

Learning Spatial-Temporal Implicit Neural Representations for Event-Guided Video Super-Resolution

2023-03-24 · CVPR 2023 1 · Yunfan Lu, Zipeng Wang, Minjie Liu, Hongjian Wang 외

Event cameras sense the intensity changes asynchronously and produce event streams with high dynamic range and low latency. This has inspired research endeavors utilizing events to guide the challenging video superresolu…

Super-ResolutionVideo Super-Resolution