paper-with-me

Papers

DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding

2026-04-28 · Cennet Oguz, Yasser Hamidullah, Josef van Genabith, Simon Ostermann arxiv

We introduce DualFact, a dual-layer, multimodal factuality evaluation framework for procedural video captioning. DualFact separates factual correctness into conceptual facts, capturing abstract semantic roles (e.g., Action, Ingredient, Tool, Location), and contextual facts, capturing their grounded predicate-argument realizations in video. To support complete and role-consistent evaluation, DualFact incorporates implicit argument augmentation (VIA) and contrastive fact sets. We instantiate DualFact in two modes: DualFact-T, which verifies facts against textual evidence, and DualFact-V, which verifies facts against video-grounded visual evidence. Experiments on YouCook3-Fact and CraftBench-Fact show that state-of-the-art multimodal language models produce fluent but often factually incomplete captions, with systematic omissions and role-level inconsistencies. DualFact correlates more strongly with human factuality judgments than standard metrics, particularly for contextual facts, and reveals that caption-only evaluation overestimates hallucinations compared to video-grounded verification. Overall, DualFact offers an interpretable and human-aligned evaluation protocol that highlights persistent challenges in multimodal factual grounding, extending beyond surface-level fluency.

📄 PDF Abstract BibTeX arXiv:2604.25584

Code (0)

등록된 구현이 없습니다.

Tasks

Fact VerificationVideo Captioning

Similar Papers 제목 키워드 기반

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

2025-09-30 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda 외 arxiv

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological exp…

Multimodal Reasoning

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning

2025-07-11 · Shibo Sun, Xue Li, Donglin Di, Mingjie Wei 외 arxiv

While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To…

How Does India Cook Biryani?

2026-01-08 · Shubham Goel, Farzana S, C V Rishi, Aditya Arun 외 arxiv

Biryani, one of India's most celebrated dishes, exhibits remarkable regional diversity in its preparation, ingredients, and presentation. With the growing availability of online cooking videos, there is unprecedented pot…

Multimodal Reasoning

Anatomy-guided Multimodal Registration by Learning Segmentation without Ground Truth: Application to Intraprocedural CBCT/MR Liver Segmentation and Registration

2021-04-14 · Bo Zhou, Zachary Augenfeld, Julius Chapiro, S. Kevin Zhou 외

Multimodal image registration has many applications in diagnostic medical imaging and image-guided interventions, such as Transcatheter Arterial Chemoembolization (TACE) of liver cancer guided by intraprocedural CBCT and…

AnatomyDiagnosticDomain AdaptationImage Registration+2

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

2025-09-03 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Susan Holm 외 arxiv

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. …