paper-with-me

Papers

Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering

2023-03-22 · Triet Minh Thai, Son T. Luu

Visual Question Answering (VQA) is a task that requires computers to give correct answers for the input questions based on the images. This task can be solved by humans with ease but is a challenge for computers. The VLSP2022-EVJVQA shared task carries the Visual Question Answering task in the multilingual domain on a newly released dataset: UIT-EVJVQA, in which the questions and answers are written in three different languages: English, Vietnamese and Japanese. We approached the challenge as a sequence-to-sequence learning task, in which we integrated hints from pre-trained state-of-the-art VQA models and image features with Convolutional Sequence-to-Sequence network to generate the desired answers. Our results obtained up to 0.3442 by F1 score on the public test set, 0.4210 on the private test set, and placed 3rd in the competition.

📄 PDF Abstract BibTeX arXiv:2303.12671

Code (1)

🤗 spaces/daeron/CONVS2S-EVJVQA-DEMO

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Unified analytic forms for Convolutional Neural Networks and Wavelet Filter Banks

2021-01-01 · Abdourrahmane M ATTO

This paper provides a unified analytic framework integrating expressions of several variants of convolutional neural networks and wavelet filter banks. The expressions are derived recursively, from dowstream to upstream …

TSCMamba: Mamba Meets Multi-View Learning for Time Series Classification

2024-06-06 · Md Atik Ahamed, Qiang Cheng

Time series classification (TSC) on multivariate time series is a critical problem. We propose a novel multi-view approach integrating frequency-domain and time-domain features to provide complementary contexts for TSC. …

MambaMULTI-VIEW LEARNINGTime SeriesTime Series Classification

Video Saliency Detection by 3D Convolutional Neural Networks

2018-07-12 · Guanqun Ding, Yuming Fang

Different from salient object detection methods for still images, a key challenging for video saliency detection is how to extract and combine spatial and temporal features. In this paper, we present a novel and effectiv…

Objectobject-detectionObject DetectionRGB Salient Object Detection+5

Integrating Mamba Sequence Model and Hierarchical Upsampling Network for Accurate Semantic Segmentation of Multiple Sclerosis Legion

2024-03-26 · Kazi Shahriar Sanjid, Md. Tanzim Hossain, Md. Shakib Shahariar Junayed, Dr. Mohammad Monir Uddin

Integrating components from convolutional neural networks and state space models in medical image segmentation presents a compelling approach to enhance accuracy and efficiency. We introduce Mamba HUNet, a novel architec…

Decision MakingImage SegmentationLesion SegmentationMamba+4

An End-to-End Khmer Optical Character Recognition using Sequence-to-Sequence with Attention

2021-06-21 · Rina Buoy, Sokchea Kor, Nguonly Taing

This paper presents an end-to-end deep convolutional recurrent neural network solution for Khmer optical character recognition (OCR) task. The proposed solution uses a sequence-to-sequence (Seq2Seq) architecture with att…

DecoderOptical Character RecognitionOptical Character Recognition (OCR)Sentence