Papers Visual Question Answering
“Visual Question Answering” 태그가 달린 논문 2,872편 · 필터 해제
Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface…
Visual Question AnsweringVDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed betw…
Visual Question AnsweringImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation
We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …
Visual Question AnsweringText-to-Image GenerationQwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D percep…
Visual Question AnsweringScene Understanding3D Object DetectionAutonomous DrivingAn Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis
An AI-based approach called ChatNDE Figure to Caption is introduced, which aims to automate the interpretation of NDE images using deep learning and natural language processing (NLP). A Vision-and-Language Pretraining (V…
Visual Question AnsweringSynth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images
The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text-reading capabilities …
Visual Question AnsweringFrom Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise pixel-level grounding…
Visual Question AnsweringMedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language…
Visual Question AnsweringAWM: Answerable Working Memory for Long-Document VQA Agents
Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers. Working memory should carry answer-s…
Visual Question AnsweringLlama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constraine…
Visual Question AnsweringMasking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI
Medical image-text data can expose protected health information (PHI) through both visible image content as well as accompanying text, creating a barrier to privacy-preserving medical AI systems. This risk is especially …
Visual Question AnsweringA Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, …
Visual Question AnsweringTrajectory PlanningAutonomous DrivingDecision MakingArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation
Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging for models to localize and recognize. Exis…
Visual Question AnsweringQuestion-Guided Evidence Acquisition for Multimodal Visual Question Answering
Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting…
Visual Question AnsweringGRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering
Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational exampl…
Visual Question AnsweringBeyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, existing retrieval pipelines predominantly employ CLIP-style dual enco…
Visual Question AnsweringPerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation
Radiology report generation has matured almost entirely on 2D chest radiographs, where the default route to better reports is a larger backbone or a pre-training one on medical data. We revisit that assumption on 3D mult…
Visual Question AnsweringCounterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insuffici…
Visual Question AnsweringSeeing Red, Thinking Bad: Color Bias in Vision Language Models
Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual informatio…
Visual Question AnsweringA Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competit…
Visual Question AnsweringInformation Extraction