paper-with-me

Papers

Multi-Task Learning of Generalizable Representations for Video Action Recognition

2018-11-20 · Zhiyu Yao, Yunbo Wang, Mingsheng Long, Jian-Min Wang, Philip S. Yu, Jiaguang Sun

In classic video action recognition, labels may not contain enough information about the diverse video appearance and dynamics, thus, existing models that are trained under the standard supervised learning paradigm may extract less generalizable features. We evaluate these models under a cross-dataset experiment setting, as the above label bias problem in video analysis is even more prominent across different data sources. We find that using the optical flows as model inputs harms the generalization ability of most video recognition models. Based on these findings, we present a multi-task learning paradigm for video classification. Our key idea is to avoid label bias and improve the generalization ability by taking data as its own supervision or supervising constraints on the data. First, we take the optical flows and the RGB frames by taking them as auxiliary supervisions, and thus naming our model as Reversed Two-Stream Networks (Rev2Net). Further, we collaborate the auxiliary flow prediction task and the frame reconstruction task by introducing a new training objective to Rev2Net, named Decoding Discrepancy Penalty (DDP), which constraints the discrepancy of the multi-task features in a self-supervised manner. Rev2Net is shown to be effective on the classic action recognition task. It specifically shows a strong generalization ability in the cross-dataset experiments.

📄 PDF Abstract BibTeX arXiv:1811.08362

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionMulti-Task LearningOptical Flow EstimationTemporal Action LocalizationVideo ClassificationVideo Recognition

Similar Papers 제목 키워드 기반

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

2026-08-31 · Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun 외 hf

Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Ego…

VidFuncta: Towards Generalizable Neural Representations for Ultrasound Videos

2025-07-29 · Julia Wolleb, Florentin Bieder, Paul Friedrich, Hemant D. Tagare 외 arxiv

Ultrasound is widely used in clinical care, yet standard deep learning methods often struggle with full video analysis due to non-standardized acquisition and operator bias. We offer a new perspective on ultrasound video…

Video ReconstructionLine Detection

Implicit State Estimation via Video Replanning

2025-10-20 · Po-Chen Ko, Jiayuan Mao, Yu-Hsiang Fu, Hsien-Jeng Yeh 외 arxiv

Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These representations enable flexible and genera…

ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping

2024-12-18 · CVPR 2025 1 · Youxin Pang, Ruizhi Shao, Jiajun Zhang, Hanzhang Tu 외

In this paper, we introduce ManiVideo, a novel method for generating consistent and temporally coherent bimanual hand-object manipulation videos from given motion sequences of hands and objects. The core idea of ManiVide…

ObjectVideo Generation

PointAction: 3D Points as Universal Action Representations for Robot Control

2026-06-02 · Mutian Tong, Han Jiang, Qiao Feng, Lingjie Liu 외 arxiv

Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only video rollouts are not di…

Robot ManipulationVideo GenerationVideo Prediction