Cascade Attention Guided Residue Learning GAN for Cross-Modal Translation
Since we were babies, we intuitively develop the ability to correlate the input from different cognitive sensors such as vision, audio, and text. However, in machine learning, this cross-modal learning is a nontrivial task because different modalities have no homogeneous properties. Previous works discover that there should be bridges among different modalities. From neurology and psychology perspective, humans have the capacity to link one modality with another one, e.g., associating a picture of a bird with the only hearing of its singing and vice versa. Is it possible for machine learning algorithms to recover the scene given the audio signal? In this paper, we propose a novel Cascade Attention-Guided Residue GAN (CAR-GAN), aiming at reconstructing the scenes given the corresponding audio signals. Particularly, we present a residue module to mitigate the gap between different modalities progressively. Moreover, a cascade attention guided network with a novel classification loss function is designed to tackle the cross-modal learning task. Our model keeps the consistency in high-level semantic label domain and is able to balance two different modalities. The experimental results demonstrate that our model achieves the state-of-the-art cross-modal audio-visual generation on the challenging Sub-URMP dataset. Code will be available at https://github.com/tuffr5/CAR-GAN.
Code (1)
Tasks
BIG-bench Machine LearningTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Attention-guided Evidence Grounding for Spoken Question Answering
Spoken Question Answering (Spoken QA) presents a challenging cross-modal problem: effectively aligning acoustic queries with textual knowledge while avoiding the latency and error propagation inherent in cascaded ASR-bas…
Question AnsweringCrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding Residues
Accurate identification of protein nucleic-acid-binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein …
Contrastive LearningDrug DesignLanguage ModellingProtein Language ModelAttentive cross-modal paratope prediction
Antibodies are a critical part of the immune system, having the function of directly neutralising or tagging undesirable objects (the antigens) for future destruction. Being able to predict which amino acids belong to th…
Antibody-antigen binding predictionComputational EfficiencyPredictionCACFNet: Cross-Modal Attention Cascaded Fusion Network for RGB-T Urban Scene Parsing
Color–thermal (RGB-T) urban scene parsing has recently attracted widespread interest. However, most existing approaches to RGB-T urban scene parsing do not deeply explore the information complementarity between RGB-T fea…
Scene ParsingThermal Image SegmentationCascaded Semantic and Positional Self-Attention Network for Document Classification
Transformers have shown great success in learning representations for language modelling. However, an open challenge still remains on how to systematically aggregate semantic information (word embedding) with positional …
ClassificationDocument ClassificationGeneral ClassificationLanguage Modelling