paper-with-me

홈 › Papers

Encoding Video and Label Priors for Multi-label Video Classification on YouTube-8M dataset

2017-06-24 · Seil Na, Youngjae Yu, Sang-ho Lee, Ji-Sung Kim, Gunhee Kim

YouTube-8M is the largest video dataset for multi-label video classification. In order to tackle the multi-label classification on this challenging dataset, it is necessary to solve several issues such as temporal modeling of videos, label imbalances, and correlations between labels. We develop a deep neural network model, which consists of four components: the frame encoder, the classification layer, the label processing layer, and the loss function. We introduce our newly proposed methods and discusses how existing models operate in the YouTube-8M Classification Task, what insights they have, and why they succeed (or fail) to achieve good performance. Most of the models we proposed are very high compared to the baseline models, and the ensemble of the models we used is 8th in the Kaggle Competition.

📄 PDF Abstract BibTeX arXiv:1706.07960

Code (1)

seilna/youtube-8m 공식 구현 tf

Tasks

ClassificationGeneral ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONVideo Classification

Similar Papers 제목 키워드 기반

Knowledge Priors for Identity-Disentangled Open-Set Privacy-Preserving Video FER

2026-03-22 · Feng Xu, Xun Li, Lars Petersson, Yulei Sui 외 arxiv

Facial expression recognition relies on facial data that inherently expose identity and thus raise significant privacy concerns. Current privacy-preserving methods typically fail in realistic open-set video settings wher…

Facial Expression Recognition

Dynamics Based Neural Encoding with Inter-Intra Region Connectivity

2024-02-19 · Mai Gamal, Mohamed Rashad, Eman Ehab, Seif Eldawlatly 외

Extensive literature has drawn comparisons between recordings of biological neurons in the brain and deep neural networks. This comparative analysis aims to advance and interpret deep neural networks and enhance our unde…

Video Understanding

Structured Label Inference for Visual Understanding

2018-02-18 · Nelson Nauata, Hexiang Hu, Guang-Tong Zhou, Zhiwei Deng 외

Visual data such as images and videos contain a rich source of structured semantic labels as well as a wide range of interacting components. Visual content could be assigned with fine-grained labels describing major comp…

Action DetectionGeneral Classificationimage-classificationImage Classification+3

Transcoded Video Restoration by Temporal Spatial Auxiliary Network

2021-12-15 · Li Xu, Gang He, Jinjia Zhou, Jie Lei 외

In most video platforms, such as Youtube, and TikTok, the played videos usually have undergone multiple video encodings such as hardware encoding by recording devices, software encoding by video editing apps, and single/…

Video EditingVideo Restoration

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

2026-05-21 · Zhanbo Huang, Xiaoming Liu, Yu Kong arxiv

Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimates dense motion but breaks down on articulated bodies, where pixel-le…