paper-with-me

홈 › Papers

FCMBench-Video: Benchmarking Document Video Intelligence

2026-04-28 · Runze Cui, Fangxin Shang, Yehui Yang, Qing Yang, Yanwu Xu, Tao Chen arxiv

Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and evidence traceability matter. Compared with static document images, document videos present a temporally redundant and sequentially unfolding evidence stream, require evidence integration across frames, and preserve acquisition-process cues relevant to authenticity-sensitive and anti-fraud review. We introduce FCMBench-Video, a benchmark for document-video intelligence that evaluates document perception, temporal grounding, and evidence-grounded reasoning under realistic capture conditions. For privacy-compliant yet realistic data at scale, we organize construction as an atomic-acquisition and composition workflow that records reusable single-document clips, applies controlled degradations, and assembles long-form multi-document videos with prescribed temporal spans. FCMBench-Video is built from 495 atomic videos composed into 1,200 long-form videos paired with 11,322 expert-annotated question--answer instances, covering 28 document types over 20s--60s duration tiers and 5,960 Chinese / 5,362 English instances. Evaluations on nine recent Video-MLLMs show that FCMBench-Video provides meaningful separation across systems and capabilities: counting is the most duration-sensitive task, Cross-Document Validation and Evidence-Grounded Selection probe higher-level evidence integration, and Visual Prompt Injection provides a complementary robustness dimension. The overall score distribution is broad and approximately bell-shaped, indicating a benchmark that is neither saturated nor dominated by trivial cases. Together, these results position FCMBench-Video as a reproducible benchmark for tracking Video-MLLM progress on document-video understanding and probing capability boundaries in authenticity-sensitive credit-domain applications.

📄 PDF Abstract BibTeX arXiv:2604.25186

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Capsule Vision 2024 Challenge: Multi-Class Abnormality Classification for Video Capsule Endoscopy

2024-08-09 · Palak Handa, Amirreza Mahbod, Florian Schwarzhans, Ramona Woitek 외

We present the Capsule Vision 2024 Challenge: Multi-Class Abnormality Classification for Video Capsule Endoscopy. It was virtually organized by the Research Center for Medical Image Analysis and Artificial Intelligence (…

BenchmarkingMedical Image Analysis

FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications

2026-01-01 · Yehui Yang, Dalu Yang, Fangxin Shang, Wenshuo Zhou 외 arxiv

FCMBench is the first large-scale and privacy-compliant multimodal benchmark for real-world financial credit applications, covering tasks and robustness challenges from domain specific workflows and constraints. The curr…

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

2026-07-17 · Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang 외 hf

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without…

Video Generation

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

2025-03-08 · Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang 외

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a benchmark to evaluate whether video-large l…

BenchmarkingcounterfactualMultiple-choice

HourVideo: 1-Hour Video-Language Understanding

2024-11-07 · Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota 외

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, tempora…

BenchmarkingcounterfactualMultiple-choiceRetrieval+1