Shuffle Transformer with Feature Alignment for Video Face Parsing
This is a short technical report introducing the solution of the Team TCParser for Short-video Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR 2021. In this paper, we introduce a strong backbone which is cross-window based Shuffle Transformer for presenting accurate face parsing representation. To further obtain the finer segmentation results, especially on the edges, we introduce a Feature Alignment Aggregation (FAA) module. It can effectively relieve the feature misalignment issue caused by multi-resolution feature aggregation. Benefiting from the stronger backbone and better feature aggregation, the proposed method achieves 86.9519% score in the Short-video Face Parsing track of the 3rd Person in Context (PIC) Workshop and Challenge, ranked the first place.
Code (0)
등록된 구현이 없습니다.
Tasks
Face ParsingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Efficient Video Representation with Video Shuffle Networks
3D CNN shows its strong ability in learning spatiotemporal representation in recent video recognition tasks. However, inflating 2D convolution to 3D inevitably introduces additional computational costs, making it cumbers…
Video RecognitionBeyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
Multiple complex degradations are coupled in low-quality video faces in the real world. Therefore, blind video face restoration is a highly challenging ill-posed problem, requiring not only hallucinating high-fidelity de…
Face ParsingSemantic ParsingVideo Temporal ConsistencyExpression Snippet Transformer for Robust Video-based Facial Expression Recognition
The recent success of Transformer has provided a new direction to various visual understanding tasks, including video-based facial expression recognition (FER). By modeling visual relations effectively, Transformer has s…
Dynamic Facial Expression RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots
In this paper, we investigate the task of hallucinating an authentic high-resolution (HR) human face from multiple low-resolution (LR) video snapshots. We propose a pure transformer-based model, dubbed VidFace, to fully …
Face HallucinationHallucinationHERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
We present HERO, a novel framework for large-scale video+language omni-representation learning. HERO encodes multimodal inputs in a hierarchical structure, where local context of a video frame is captured by a Cross-moda…
Language ModelingLanguage ModellingMasked Language ModelingMoment Retrieval+7