paper-with-me

Papers

Cooperative Cross-Stream Network for Discriminative Action Representation

2019-08-27 · Jingran Zhang, Fumin Shen, Xing Xu, Heng Tao Shen

Spatial and temporal stream model has gained great success in video action recognition. Most existing works pay more attention to designing effective features fusion methods, which train the two-stream model in a separate way. However, it's hard to ensure discriminability and explore complementary information between different streams in existing works. In this work, we propose a novel cooperative cross-stream network that investigates the conjoint information in multiple different modalities. The jointly spatial and temporal stream networks feature extraction is accomplished by an end-to-end learning manner. It extracts this complementary information of different modality from a connection block, which aims at exploring correlations of different stream features. Furthermore, different from the conventional ConvNet that learns the deep separable features with only one cross-entropy loss, our proposed model enhances the discriminative power of the deeply learned features and reduces the undesired modality discrepancy by jointly optimizing a modality ranking constraint and a cross-entropy loss for both homogeneous and heterogeneous modalities. The modality ranking constraint constitutes intra-modality discriminative embedding and inter-modality triplet constraint, and it reduces both the intra-modality and cross-modality feature variations. Experiments on three benchmark datasets demonstrate that by cooperating appearance and motion feature extraction, our method can achieve state-of-the-art or competitive performance compared with existing results.

📄 PDF Abstract BibTeX arXiv:1908.10136

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action LocalizationTriplet

Similar Papers 제목 키워드 기반

QUEST: Query Stream for Practical Cooperative Perception

2023-08-03 · Siqi Fan, Haibao Yu, Wenxian Yang, Jirui Yuan 외

Cooperative perception can effectively enhance individual perception performance by providing additional viewpoint and expanding the sensing field. Existing cooperation paradigms are either interpretable (result cooperat…

3D Object Detection

Cross-Model Cross-Stream Learning for Self-Supervised Human Action Recognition

2023-07-15 · Mengyuan Liu, Hong Liu, Tianyu Guo

Considering the instance-level discriminative ability, contrastive learning methods, including MoCo and SimCLR, have been adapted from the original image representation learning task to solve the self-supervised skeleton…

Action RecognitionContrastive LearningEnsemble LearningPseudo Label+6

Cross-Model Cross-Stream Learning for Self-Supervised Human Action Recognition

2024-09-23 · IEEE Transactions on Human-Machine Systems 2024 9 · Liu, Mengyuan; Liu, Hong; Guo, Tianyu

Considering the instance-level discriminative ability, contrastive learning methods, including MoCo and SimCLR, have been adapted from the original image representation learning task to solve the self-supervised skeleton…

Action RecognitionContrastive LearningEnsemble LearningPseudo Label+5

Enhancing Text Generation with Cooperative Training

2023-03-16 · Tong Wu, Hao Wang, Zhongshen Zeng, Wei Wang 외

Recently, there has been a surge in the use of generated data to enhance the performance of downstream models, largely due to the advancements in pre-trained language models. However, most prevailing methods trained gene…

MRPCQQPSTSText Generation

Multi-scale Cooperative Multimodal Transformers for Multimodal Sentiment Analysis in Videos

2022-06-16 · Lianyang Ma, Yu Yao, Tao Liang, Tongliang Liu

Multimodal sentiment analysis in videos is a key task in many real-world applications, which usually requires integrating multimodal streams including visual, verbal and acoustic behaviors. To improve the robustness of m…

Multimodal Sentiment AnalysisSentiment Analysis