paper-with-me

홈 › Papers

AttnGrounder: Talking to Cars with Attention

2020-09-11 · Vivek Mittal

We propose Attention Grounder (AttnGrounder), a single-stage end-to-end trainable model for the task of visual grounding. Visual grounding aims to localize a specific object in an image based on a given natural language text query. Unlike previous methods that use the same text representation for every image region, we use a visual-text attention module that relates each word in the given query with every region in the corresponding image for constructing a region dependent text representation. Furthermore, for improving the localization ability of our model, we use our visual-text attention module to generate an attention mask around the referred object. The attention mask is trained as an auxiliary task using a rectangular mask generated with the provided ground-truth coordinates. We evaluate AttnGrounder on the Talk2Car dataset and show an improvement of 3.26% over the existing methods.

📄 PDF Abstract BibTeX arXiv:2009.05684

Code (1)

i-m-vivek/AttnGrounder 공식 구현 pytorch

Tasks

Referring Expression ComprehensionVisual Grounding

Similar Papers 제목 키워드 기반

Talking-Heads Attention

2020-03-05 · Noam Shazeer, Zhenzhong Lan, Youlong Cheng, Nan Ding 외

We introduce "talking-heads attention" - a variation on multi-head attention which includes linearprojections across the attention-heads dimension, immediately before and after the softmax operation.While inserting only …

Language ModelingLanguage ModellingMasked Language ModelingQuestion Answering+1

An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

2022-03-10 · Ganglai Wang, Peng Zhang, Lei Xie, Wei Huang 외

DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake video detection is further improved. By on…

Decision MakingFace DetectionFace GenerationFace Swapping+1

Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented Networks

2023-06-06 · Jianrong Wang, Yaxin Zhao, Li Liu, Tianyi Xu 외

Given an audio clip and a reference face image, the goal of the talking head generation is to generate a high-fidelity talking head video. Although some audio-driven methods of generating talking head videos have made so…

Talking Head Generation

Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild

2022-03-08 · Ganglai Wang, Peng Zhang, Lei Xie, Wei Huang 외

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. …

Face GenerationTalking Face Generation

FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

2023-07-08 · Ganglai Wang, Peng Zhang, Junwen Xiong, Feihan Yang 외

DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake video detection is further improved. By on…

Face DetectionFace GenerationFace SwappingOptical Flow Estimation+1