paper-with-me

홈 › Papers

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

2026-05-08 · Hengyi Feng, Hao Liang, Mingrui Chen, Bohan Zeng, Meiyi Qiang, Zhengyang Zhao, Zimo Meng, Zeang Sheng, Wentao Zhang arxiv

Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capability. They restrict videos to short clips, isolate modalities, or reduce questions to one-hop perception. We introduce TraceAV-Bench, the first benchmark to jointly evaluate multi-hop reasoning over long audio-visual trajectories and multimodal hallucination robustness. TraceAV-Bench comprises 2,200 rigorously validated multiple-choice questions over 578 long videos, totaling 339.5 hours, spanning 4 evaluation dimensions and 15 sub-tasks. Each question is grounded in an explicit reasoning chain that averages 3.68 hops across a 15.1-minute temporal span. The dataset is built by a three-step semi-automated pipeline followed by a strict quality assurance process. Evaluation of multiple representative OmniLLMs on TraceAV-Bench reveals that the benchmark poses a persistent challenge across all models, with the strongest closed-source model (Gemini 3.1 Pro) reaching only 68.29% on general tasks, and the best open-source model (Ming-Flash-Omni-2.0) reaching 51.70%, leaving substantial headroom. Moreover, we find that robustness to multimodal hallucination is largely decoupled from general multimodal reasoning performance. We anticipate that TraceAV-Bench will stimulate further research toward OmniLLMs that can reason coherently and faithfully over long-form audio-visual content.

📄 PDF Abstract BibTeX arXiv:2605.07593

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Quantifying the Complexity of Standard Benchmarking Datasets for Long-Term Human Trajectory Prediction

2020-05-28 · Ronny Hug, Stefan Becker, Wolfgang Hübner, Michael Arens

Methods to quantify the complexity of trajectory datasets are still a missing piece in benchmarking human trajectory prediction models. In order to gain a better understanding of the complexity of trajectory prediction t…

BenchmarkingPredictionQuantizationTrajectory Prediction

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

2026-05-09 · Xiang Feng, Jiawei Zhou, Zhangfeng Huang, Kewei Wang 외 arxiv

Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer accuracy. We introduce DocScope, a benchma…

Trajectory Prediction

Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding

2025-09-30 · Haotian Xue, Yunhao Ge, Yu Zeng, Zhaoshuo Li 외 arxiv

Vision-Language Models (VLMs) have demonstrated impressive world knowledge across a wide range of tasks, making them promising candidates for embodied reasoning applications. However, existing benchmarks primarily evalua…

Object LocalizationVisual Grounding

PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

2026-06-03 · Hainiu Xu, Italo Luis da Silva, Jiangnan Ye, Yuhao Wang 외 arxiv

Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective. While existing safety research has focused on detecting unethical …

OpenTraj: Assessing Prediction Complexity in Human Trajectories Datasets

2020-10-02 · Javad Amirian, Bingqing Zhang, Francisco Valente Castro, Juan Jose Baldelomar 외

Human Trajectory Prediction (HTP) has gained much momentum in the last years and many solutions have been proposed to solve it. Proper benchmarking being a key issue for comparing methods, this paper addresses the questi…

BenchmarkingPredictionSelf-Driving CarsTrajectory Forecasting+1