paper-with-me

Papers

MCQA: Multimodal Co-attention Based Network for Question Answering

2020-04-25 · Abhishek Kumar, Trisha Mittal, Dinesh Manocha

We present MCQA, a learning-based algorithm for multimodal question answering. MCQA explicitly fuses and aligns the multimodal input (i.e. text, audio, and video), which forms the context for the query (question and answer). Our approach fuses and aligns the question and the answer within this context. Moreover, we use the notion of co-attention to perform cross-modal alignment and multimodal context-query alignment. Our context-query alignment module matches the relevant parts of the multimodal context and the query with each other and aligns them to improve the overall performance. We evaluate the performance of MCQA on Social-IQ, a benchmark dataset for multimodal question answering. We compare the performance of our algorithm with prior methods and observe an accuracy improvement of 4-7%.

📄 PDF Abstract BibTeX arXiv:2004.12238

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentQuestion Answering

Similar Papers 제목 키워드 기반

Afri-MCQA: Multimodal Cultural Question Answering for African Languages

2026-01-09 · Atnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime 외 arxiv

Africa is home to over one-third of the world's languages, yet remains underrepresented in AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering benchmark covering 7.5k Q&A pairs across …

Question Answering

Differentiating Choices via Commonality for Multiple-Choice Question Answering

2024-08-21 · Wenqing Deng, Zhe Wang, Kewen Wang, Shirui Pan 외

Multiple-choice question answering (MCQA) becomes particularly challenging when all choices are relevant to the question and are semantically similar. Yet this setting of MCQA can potentially provide valuable clues for c…

Multiple-choiceMultiple Choice Question Answering (MCQA)Question Answering

KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination

2026-02-14 · Byungjin Choi, Seongsu Bae, Sunjun Kweon, Edward Choi arxiv

We introduce KorMedMCQA-V, a Korean medical licensing-exam-style multimodal multiple-choice question answering benchmark for evaluating vision-language models (VLMs). The dataset consists of 1,534 questions with 2,043 as…

Question Answering

FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain

2023-04-09 · LOUHI 2022 10 · Yanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier 외

This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain. It is composed of 3,105 questions taken from real exams of the French medi…

Multiple-choiceMultiple Choice Question Answering (MCQA)Question Answering

Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT

2025-11-21 · Yesheng Liu, Hao Li, Haiyu Xu, Baoqi Pei 외 arxiv

Multiple-choice question answering (MCQA) has been a popular format for evaluating and reinforcement fine-tuning (RFT) of modern multimodal language models. Its constrained output format allows for simplified, determinis…

Question Answering