paper-with-me

홈 › Papers

VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering

2026-02-03 · Rahul Atul Bhope, K. R. Jayaram, Vinod Muthusamy, Ritesh Kumar, Vatche Isahagian, Nalini Venkatasubramanian arxiv

Despite significant costs from retrieving and processing high-fidelity visual inputs, most multimodal vision-language systems operate at fixed fidelity levels. We introduce VOILA, a framework for Value-Of-Information-driven adaptive fidelity selection in Visual Question Answering (VQA) that optimizes what information to retrieve before model execution. Given a query, VOILA uses a two-stage pipeline: a gradient-boosted regressor estimates correctness likelihood at each fidelity from question features alone, then an isotonic calibrator refines these probabilities for reliable decision-making. The system selects the minimum-cost fidelity maximizing expected utility given predicted accuracy and retrieval costs. We evaluate VOILA across three deployment scenarios using five datasets (VQA-v2, GQA, TextVQA, LoCoMo, FloodNet) and six Vision-Language Models (VLMs) with 7B-235B parameters. VOILA consistently achieves 50-60% cost reductions while retaining 90-95% of full-resolution accuracy across diverse query types and model architectures, demonstrating that pre-retrieval fidelity selection is vital to optimize multimodal inference under resource constraints.

📄 PDF Abstract BibTeX arXiv:2602.03007

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Value of Information Lattice: Exploiting Probabilistic Independence for Effective Feature Subset Acquisition

2014-01-16 · Mustafa Bilgic, Lise Getoor

We address the cost-sensitive feature acquisition problem, where misclassifying an instance is costly but the expected misclassification cost can be reduced by acquiring the values of the missing features. Because acquir…

Voila-A: Aligning Vision-Language Models with User's Gaze Attention

2023-12-22 · Kun Yan, Lei Ji, Zeyu Wang, Yuntao Wang 외

In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challe…

G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios

2024-05-13 · Zeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao 외

Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze -- a modality deeply linked to user intent and increasingly accessible via gaz…

Natural Language Queries

VOILA: An Optimised Dialogue System for Interactively Learning Visually-Grounded Word Meanings (Demonstration System)

2017-08-01 · WS 2017 8 · Yanchao Yu, Arash Eshghi, Oliver Lemon

We present VOILA: an optimised, multi-modal dialogue agent for interactive learning of visually grounded word meanings from a human user. VOILA is: (1) able to learn new visual categories interactively from users from sc…

Active Learning

VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents

2026-06-18 · Marcus Hoerger, Rishikesh Joshi, Rahul Shome, Ian Manchester 외 arxiv

Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provides a powerful framework for such a capability. Although POMDP-based planning has…