paper-with-me

Papers

Cascade Attention Guided Residue Learning GAN for Cross-Modal Translation

2019-07-03 · Bin Duan, Wei Wang, Hao Tang, Hugo Latapie, Yan Yan

Since we were babies, we intuitively develop the ability to correlate the input from different cognitive sensors such as vision, audio, and text. However, in machine learning, this cross-modal learning is a nontrivial task because different modalities have no homogeneous properties. Previous works discover that there should be bridges among different modalities. From neurology and psychology perspective, humans have the capacity to link one modality with another one, e.g., associating a picture of a bird with the only hearing of its singing and vice versa. Is it possible for machine learning algorithms to recover the scene given the audio signal? In this paper, we propose a novel Cascade Attention-Guided Residue GAN (CAR-GAN), aiming at reconstructing the scenes given the corresponding audio signals. Particularly, we present a residue module to mitigate the gap between different modalities progressively. Moreover, a cascade attention guided network with a novel classification loss function is designed to tackle the cross-modal learning task. Our model keeps the consistency in high-level semantic label domain and is able to balance two different modalities. The experimental results demonstrate that our model achieves the state-of-the-art cross-modal audio-visual generation on the challenging Sub-URMP dataset. Code will be available at https://github.com/tuffr5/CAR-GAN.

📄 PDF Abstract BibTeX arXiv:1907.01826

Code (1)

tuffr5/CAR-GAN 공식 구현 pytorch

Tasks

BIG-bench Machine LearningTranslation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Attention-guided Evidence Grounding for Spoken Question Answering

2026-03-17 · Ke Yang, Bolin Chen, Yuejie Li, Yueying Hua 외 arxiv

Spoken Question Answering (Spoken QA) presents a challenging cross-modal problem: effectively aligning acoustic queries with textual knowledge while avoiding the latency and error propagation inherent in cascaded ASR-bas…

Question Answering

CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding Residues

2023-12-19 · Linglin Jing, Sheng Xu, Yifan Wang, Yuzhe Zhou 외

Accurate identification of protein nucleic-acid-binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein …

Contrastive LearningDrug DesignLanguage ModellingProtein Language Model

Attentive cross-modal paratope prediction

2018-06-12 · Andreea Deac, Petar Veličković, Pietro Sormanni

Antibodies are a critical part of the immune system, having the function of directly neutralising or tagging undesirable objects (the antigens) for future destruction. Being able to predict which amino acids belong to th…

Antibody-antigen binding predictionComputational EfficiencyPrediction

CACFNet: Cross-Modal Attention Cascaded Fusion Network for RGB-T Urban Scene Parsing

2023-09-14 · journal 2023 9 · WuJie Zhou, Shaohua Dong, Meixin Fang, Lu Yu

Color–thermal (RGB-T) urban scene parsing has recently attracted widespread interest. However, most existing approaches to RGB-T urban scene parsing do not deeply explore the information complementarity between RGB-T fea…

Scene ParsingThermal Image Segmentation

Cascaded Semantic and Positional Self-Attention Network for Document Classification

2020-09-15 · Findings of the Association for Computational Linguistics 2020 · Juyong Jiang, Jie Zhang, Kai Zhang

Transformers have shown great success in learning representations for language modelling. However, an open challenge still remains on how to systematically aggregate semantic information (word embedding) with positional …

ClassificationDocument ClassificationGeneral ClassificationLanguage Modelling