paper-with-me

홈 › Papers

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

2026-07-29 · Feixiang Liu, Qiang Qiu, Hao Zhang, Xinyue Wang arxiv

Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.

📄 PDF Abstract BibTeX arXiv:2608.00077

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Data Provenance via Differential Auditing

2022-09-04 · Xin Mu, Ming Pang, Feida Zhu

Auditing Data Provenance (ADP), i.e., auditing if a certain piece of data has been used to train a machine learning model, is an important problem in data provenance. The feasibility of the task has been demonstrated by …

AI Supply Chain Galaxy: 3D Visual Analytics for License Compliance

2026-06-15 · Weiru Han, Xuetao Shi, Wenyi He, Wei Wang 외 arxiv

The rapid proliferation of machine learning model reuse has transformed the AI ecosystem into a highly interconnected supply chain. Traditional compliance tools and static reports struggle to navigate these massive, mult…

Community Detection

Provable Model Provenance Set for Large Language Models

2026-01-31 · Xiaoqi Qiu, Hao Zeng, Zhiyu Hou, Hongxin Wei arxiv

The growing prevalence of unauthorized model usage and misattribution has increased the need for reliable model provenance analysis. However, existing methods largely rely on heuristic fingerprint-matching rules that lac…

MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents

2025-08-29 · Xijia Tao, Yihua Teng, Xinxing Su, Xinyu Fu 외 arxiv

Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verification. We introduce MMSearch-Plus, a 311…

Multimodal ReasoningText Retrieval

DeepProv: Behavioral Characterization and Repair of Neural Networks via Inference Provenance Graph Analysis

2025-09-30 · Firas Ben Hmida, Abderrahmen Amich, Ata Kaboudi, Birhanu Eshete arxiv

Deep neural networks (DNNs) are increasingly being deployed in high-stakes applications, from self-driving cars to biometric authentication. However, their unpredictable and unreliable behaviors in real-world settings re…

Adversarial Robustness