paper-with-me

Papers

AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and Localization

2019-11-27 · Xiao-Yu Zhang, Changsheng Li, Haichao Shi, Xiaobin Zhu, Peng Li, Jing Dong

The point process is a solid framework to model sequential data, such as videos, by exploring the underlying relevance. As a challenging problem for high-level video understanding, weakly supervised action recognition and localization in untrimmed videos has attracted intensive research attention. Knowledge transfer by leveraging the publicly available trimmed videos as external guidance is a promising attempt to make up for the coarse-grained video-level annotation and improve the generalization performance. However, unconstrained knowledge transfer may bring about irrelevant noise and jeopardize the learning model. This paper proposes a novel adaptability decomposing encoder-decoder network to transfer reliable knowledge between trimmed and untrimmed videos for action recognition and localization via bidirectional point process modeling, given only video-level annotations. By decomposing the original features into domain-adaptable and domain-specific ones based on their adaptability, trimmed-untrimmed knowledge transfer can be safely confined within a more coherent subspace. An encoder-decoder based structure is carefully designed and jointly optimized to facilitate effective action classification and temporal localization. Extensive experiments are conducted on two benchmark datasets (i.e., THUMOS14 and ActivityNet1.3), and experimental results clearly corroborate the efficacy of our method.

📄 PDF Abstract BibTeX arXiv:1911.11961

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction RecognitionDecoderTemporal LocalizationTransfer LearningVideo UnderstandingWeakly-Supervised Action Recognition

Similar Papers 제목 키워드 기반

Training ASR models by Generation of Contextual Information

2019-10-27 · Kritika Singh, Dmytro Okhonko, Jun Liu, Yongqiang Wang 외

Supervised ASR models have reached unprecedented levels of accuracy, thanks in part to ever-increasing amounts of labelled training data. However, in many applications and locales, only moderate amounts of data are avail…

Decoderspeech-recognitionSpeech RecognitionText Generation+1

Content Adaptive Latents and Decoder for Neural Image Compression

2022-12-20 · Guanbo Pan, Guo Lu, Zhihao Hu, Dong Xu

In recent years, neural image compression (NIC) algorithms have shown powerful coding performance. However, most of them are not adaptive to the image content. Although several content adaptive methods have been proposed…

DecoderImage Compression

Self-Supervised Model Adaptation for Multimodal Semantic Segmentation

2018-08-11 · Abhinav Valada, Rohit Mohan, Wolfram Burgard

Learning to reliably perceive and understand the scene is an integral enabler for robots to operate in the real-world. This problem is inherently challenging due to the multitude of object types as well as appearance cha…

DecodermodelScene RecognitionSemantic Segmentation

AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval

2024-05-28 · Sihe Zhang, Qingdong He, Jinlong Peng, Yuxi Li 외

Image retrieval aims to identify visually similar images within a database using a given query image. Traditional methods typically employ both global and local features extracted from images for matching, and may also a…

Image RetrievalRe-RankingRetrieval

Large scale weakly and semi-supervised learning for low-resource video ASR

2020-05-16 · Kritika Singh, Vimal Manohar, Alex Xiao, Sergey Edunov 외

Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of transcribing social media videos in low-…

Decoderspeech-recognitionSpeech Recognition