paper-with-me

Papers

Improving Action Localization by Progressive Cross-stream Cooperation

2019-05-28 · CVPR 2019 6 · Rui Su, Wanli Ouyang, Luping Zhou, Dong Xu

Spatio-temporal action localization consists of three levels of tasks: spatial localization, action classification, and temporal segmentation. In this work, we propose a new Progressive Cross-stream Cooperation (PCSC) framework to use both region proposals and features from one stream (i.e. Flow/RGB) to help another stream (i.e. RGB/Flow) to iteratively improve action localization results and generate better bounding boxes in an iterative fashion. Specifically, we first generate a larger set of region proposals by combining the latest region proposals from both streams, from which we can readily obtain a larger set of labelled training samples to help learn better action detection models. Second, we also propose a new message passing approach to pass information from one stream to another stream in order to learn better representations, which also leads to better action detection models. As a result, our iterative framework progressively improves action localization results at the frame level. To improve action localization results at the video level, we additionally propose a new strategy to train class-specific actionness detectors for better temporal segmentation, which can be readily learnt by focusing on "confusing" samples from the same action class. Comprehensive experiments on two benchmark datasets UCF-101-24 and J-HMDB demonstrate the effectiveness of our newly proposed approaches for spatio-temporal action localization in realistic scenarios.

📄 PDF Abstract BibTeX arXiv:1905.11575

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction DetectionAction LocalizationSpatio-Temporal Action LocalizationTemporal Action Localization

Similar Papers 제목 키워드 기반

QUEST: Query Stream for Practical Cooperative Perception

2023-08-03 · Siqi Fan, Haibao Yu, Wenxian Yang, Jirui Yuan 외

Cooperative perception can effectively enhance individual perception performance by providing additional viewpoint and expanding the sensing field. Existing cooperation paradigms are either interpretable (result cooperat…

3D Object Detection

Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

2025-03-17 · CVPR 2025 1 · Henghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang 외

In recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal localization, spatial localization, spatio-te…

Data InteractionScene UnderstandingTemporal LocalizationUIE

Rethinking Air-Ground Collaboration: A Progressive Cross-Task Benchmark and Socialized Learning Framework

2026-06-17 · Zhoupeng Guo, Yunqi Zhu, Zhihe Fan, Xinjie Yao 외 arxiv

Air-ground collaborative perception is crucial for robust visual understanding in real-world dynamic environments. However, existing studies typically formulate collaboration as single-task cross-view fusion, overlooking…

What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control

2026-04-29 · Paraskevas V. Lekeas, Giorgos Stamatopoulos arxiv

LLM agents are known to deviate from Nash equilibria in strategic interactions, but nobody has looked inside the model to understand why, or asked whether the deviation can be reversed. We do both. Working with four open…

Anchor-free temporal action localization via Progressive Boundary-aware Boosting

2023-01-01 · journal 2023 1 · Yepeng Tang, Weining Wang, Yanwu Yang, Chunjie Zhang 외

Enormous untrimmed videos from the real world are difficult to analyze and manage. Temporal action localization algorithms can help us to locate and recognize human activity clips in untrimmed videos. Recently, anchor-fr…

Action LocalizationTemporal Action Localization