paper-with-me

Papers

Aligned Dual Channel Graph Convolutional Network for Visual Question Answering

2020-07-01 · ACL 2020 6 · Qingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng, Junying Chen, Ho-fung Leung, Qing Li

Visual question answering aims to answer the natural language question about a given image. Existing graph-based methods only focus on the relations between objects in an image and neglect the importance of the syntactic dependency relations between words in a question. To simultaneously capture the relations between objects in an image and the syntactic dependency relations between words in a question, we propose a novel dual channel graph convolutional network (DC-GCN) for better combining visual and textual advantages. The DC-GCN model consists of three parts: an I-GCN module to capture the relations between objects in an image, a Q-GCN module to capture the syntactic dependency relations between words in a question, and an attention alignment module to align image representations and question representations. Experimental results show that our model achieves comparable performance with the state-of-the-art approaches.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

DBConformer: Dual-Branch Convolutional Transformer for EEG Decoding

2025-06-26 · Ziwei Wang, Hongbin Wang, Tianwang Jia, Xingyi He 외

Electroencephalography (EEG)-based brain-computer interfaces (BCIs) transform spontaneous/evoked neural activity into control commands for external communication. While convolutional neural networks (CNNs) remain the mai…

EEGEeg DecodingMotor ImagerySeizure Detection

Dual Stream Computer-Generated Image Detection Network Based On Channel Joint And Softpool

2022-07-07 · Ziyi Xi, Hao Lin, Weiqi Luo

With the development of computer graphics technology, the images synthesized by computer software become more and more closer to the photographs. While computer graphics technology brings us a grand visual feast in the f…

Image Forensics

Select and Calibrate the Low-confidence: Dual-Channel Consistency based Graph Convolutional Networks

2022-05-08 · Shuhao Shi, Jian Chen, Kai Qiao, Shuai Yang 외

The Graph Convolutional Networks (GCNs) have achieved excellent results in node classification tasks, but the model's performance at low label rates is still unsatisfactory. Previous studies in Semi-Supervised Learning (…

Node Classification

Dissociating spatial frequency reliance from adversarial robustness advantages in neurally guided deep convolutional neural networks

2026-05-06 · Zhenan Shao, Tianyu Ren, Chengxiao Wang, Leyla Isik 외 arxiv

Deep convolutional neural networks (DCNNs) have rivaled humans on many visual tasks, yet they remain vulnerable to near-imperceptible perturbations generated by adversarial attacks. Recent work shows that aligning DCNN r…

Adversarial RobustnessObject Recognition

Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations

2026-02-06 · Shamik Bhattacharya, Daniel Perkins, Yaren Dogan, Vineeth Konjeti 외 arxiv

Ambiguity poses persistent challenges in natural language understanding for large language models (LLMs). To better understand how lexical ambiguity can be resolved through the visual domain, we develop an interpretable …

Natural Language UnderstandingWord Sense DisambiguationImage Augmentation