paper-with-me

Papers Visual Question Answering

“Visual Question Answering” 태그가 달린 논문 2,872편 · 필터 해제

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

2026-09-09 · Mingbo Yang, Wenqiang Wang, Zhaolu Kang, Peng Chen 외 arxiv

In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface…

Visual Question Answering

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

2026-09-05 · Yixin Wan, Tianle Zheng, Kai-Wei Chang hf

Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed betw…

Visual Question Answering

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

2026-08-31 · Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir 외 arxiv

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …

Visual Question AnsweringText-to-Image Generation

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

2026-08-31 · Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li 외 hf

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D percep…

Visual Question AnsweringScene Understanding3D Object DetectionAutonomous Driving

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

2026-08-29 · Mehrdad Shafiei Dizaji, Hoda Azari arxiv

An AI-based approach called ChatNDE Figure to Caption is introduced, which aims to automate the interpretation of NDE images using deep learning and natural language processing (NLP). A Vision-and-Language Pretraining (V…

Visual Question Answering

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

2026-08-28 · Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara arxiv

The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text-reading capabilities …

Visual Question Answering

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

2026-08-27 · Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang 외 arxiv

Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise pixel-level grounding…

Visual Question Answering

MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

2026-08-27 · Haowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren 외 arxiv

Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language…

Visual Question Answering

AWM: Answerable Working Memory for Long-Document VQA Agents

2026-08-26 · Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu, Zhen Yang 외 arxiv

Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers. Working memory should carry answer-s…

Visual Question Answering

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

2026-08-21 · Luka Ribar, Jeevan Bhoot, Douglas Orr arxiv

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constraine…

Visual Question Answering

Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI

2026-08-21 · Shiva Shrestha, Zongxing Xie, Chen Zhao, Liran Ma 외 arxiv

Medical image-text data can expose protected health information (PHI) through both visible image content as well as accompanying text, creating a barrier to privacy-preserving medical AI systems. This risk is especially …

Visual Question Answering

A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving

2026-08-21 · Jingtao Sun, Xiaohai He, Yike Zhang, Dong Huang 외 arxiv

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, …

Visual Question AnsweringTrajectory PlanningAutonomous DrivingDecision Making

ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

2026-08-20 · Linhan Cao, Siyuan Li, Jun Lan, Liangbo He 외 arxiv

Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging for models to localize and recognize. Exis…

Visual Question Answering

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

2026-08-20 · Alin-Ionut Popa arxiv

Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting…

Visual Question Answering

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

2026-08-19 · Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu 외 arxiv

Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational exampl…

Visual Question Answering

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

2026-08-19 · Hangrui Xu, Zhengxian Wu, Yunyao Yu, Zhuohong Chen 외 hf

Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, existing retrieval pipelines predominantly employ CLIP-style dual enco…

Visual Question Answering

PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation

2026-08-18 · Jianyu Sun, Zhenxuan Zhang, Guang Yang, Peter J. Lally arxiv

Radiology report generation has matured almost entirely on 2D chest radiographs, where the default route to better reports is a larger backbone or a pre-training one on medical data. We revisit that assumption on 3D mult…

Visual Question Answering

Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs

2026-08-18 · Yifan Lu, Adinath Dukre, Abhijit Das, Ziyun Zou 외 arxiv

Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insuffici…

Visual Question Answering

Seeing Red, Thinking Bad: Color Bias in Vision Language Models

2026-08-14 · Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka 외 arxiv

Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual informatio…

Visual Question Answering

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

2026-08-14 · Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova 외 arxiv

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competit…

Visual Question AnsweringInformation Extraction
1–20 / 2,872 다음 →