paper-with-me

Papers

Technical Report: Disentangled Action Parsing Networks for Accurate Part-level Action Parsing

2021-11-05 · Xuanhan Wang, Xiaojia Chen, Lianli Gao, Lechao Chen, Jingkuan Song

Part-level Action Parsing aims at part state parsing for boosting action recognition in videos. Despite of dramatic progresses in the area of video classification research, a severe problem faced by the community is that the detailed understanding of human actions is ignored. Our motivation is that parsing human actions needs to build models that focus on the specific problem. We present a simple yet effective approach, named disentangled action parsing (DAP). Specifically, we divided the part-level action parsing into three stages: 1) person detection, where a person detector is adopted to detect all persons from videos as well as performs instance-level action recognition; 2) Part parsing, where a part-parsing model is proposed to recognize human parts from detected person images; and 3) Action parsing, where a multi-modal action parsing network is used to parse action category conditioning on all detection results that are obtained from previous stages. With these three major models applied, our approach of DAP records a global mean of $0.605$ score in 2021 Kinetics-TPS Challenge.

📄 PDF Abstract BibTeX arXiv:2111.03225

Code (0)

등록된 구현이 없습니다.

Tasks

Action ParsingAction RecognitionAction Recognition In VideosHuman DetectionVideo Classification

Similar Papers 제목 키워드 기반

A Baseline Framework for Part-level Action Parsing and Action Recognition

2021-10-07 · Xiaodong Chen, Xinchen Liu, Kun Liu, Wu Liu 외

This technical report introduces our 2nd place solution to Kinetics-TPS Track on Part-level Action Parsing in ICCV DeeperAction Workshop 2021. Our entry is mainly based on YOLOF for instance and part detection, HRNet for…

Action ParsingAction RecognitionPose Estimation

PaddleOCR 3.0 Technical Report

2025-07-08 · Cheng Cui, Ting Sun, Manhui Lin, Tingquan Gao 외

This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the era of large language models, PaddleOCR…

document understandingKey Information ExtractionOptical Character Recognition (OCR)

Uni-Parser Technical Report

2025-12-17 · Xi Fang, Haoyi Tao, Shuwen Yang, Chaozheng Huang 외 arxiv

This technical report introduces Uni-Parser, an industrial-grade document parsing engine tailored for scientific literature and patents, delivering high throughput, robust accuracy, and cost efficiency. Unlike pipeline-b…

Shuffle Transformer with Feature Alignment for Video Face Parsing

2021-06-16 · Rui Zhang, Yang Han, Zilong Huang, Pei Cheng 외

This is a short technical report introducing the solution of the Team TCParser for Short-video Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR 2021. In this paper, we introduce a stro…

Face Parsing

3rd Place Solution for Short-video Face Parsing Challenge

2021-06-14 · Xiao Liu, XiaoFei Si, Jiangtao Xie

This is a short technical report introducing the solution of Team Rat for Short-video Parsing Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR 2021. In this report, we propose an Edge-…

Face Parsing