paper-with-me

Papers

GGViT:Multistream Vision Transformer Network in Face2Face Facial Reenactment Detection

2022-10-12 · Haotian Wu, Peipei Wang, Xin Wang, Ji Xiang, Rui Gong

Detecting manipulated facial images and videos on social networks has been an urgent problem to be solved. The compression of videos on social media has destroyed some pixel details that could be used to detect forgeries. Hence, it is crucial to detect manipulated faces in videos of different quality. We propose a new multi-stream network architecture named GGViT, which utilizes global information to improve the generalization of the model. The embedding of the whole face extracted by ViT will guide each stream network. Through a large number of experiments, we have proved that our proposed model achieves state-of-the-art classification accuracy on FF++ dataset, and has been greatly improved on scenarios of different compression rates. The accuracy of Raw/C23, Raw/C40 and C23/C40 was increased by 24.34%, 15.08% and 10.14% respectively.

📄 PDF Abstract BibTeX arXiv:2210.05990

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multistream CNN for Robust Acoustic Modeling

2020-05-21 · Kyu J. Han, Jing Pan, Venkata Krishna Naveen Tadala, Tao Ma 외

This paper proposes multistream CNN, a novel neural network architecture for robust acoustic modeling in speech recognition tasks. The proposed architecture processes input speech with diverse temporal resolutions by app…

Data Augmentationspeech-recognitionSpeech Recognition

Part-based Face Recognition with Vision Transformers

2022-11-30 · Zhonglin Sun, Georgios Tzimiropoulos

Holistic methods using CNNs and margin-based losses have dominated research on face recognition. In this work, we depart from this setting in two ways: (a) we employ the Vision Transformer as an architecture for training…

Face Recognition

Surface Analysis with Vision Transformers

2022-05-31 · Simon Dahan, Logan Z. J. Williams, Abdulah Fawaz, Daniel Rueckert 외

The extension of convolutional neural networks (CNNs) to non-Euclidean geometries has led to multiple frameworks for studying manifolds. Many of those methods have shown design limitations resulting in poor modelling of …

Surface Vision Transformers: Flexible Attention-Based Modelling of Biomedical Surfaces

2022-04-07 · Simon Dahan, Hao Xu, Logan Z. J. Williams, Abdulah Fawaz 외

Recent state-of-the-art performances of Vision Transformers (ViT) in computer vision tasks demonstrate that a general-purpose architecture, which implements long-range self-attention, could replace the local feature lear…

ClassificationData Augmentation

Surface Vision Transformers: Attention-Based Modelling applied to Cortical Analysis

2022-03-30 · Simon Dahan, Abdulah Fawaz, Logan Z. J. Williams, Chunhui Yang 외

The extension of convolutional neural networks (CNNs) to non-Euclidean geometries has led to multiple frameworks for studying manifolds. Many of those methods have shown design limitations resulting in poor modelling of …