paper-with-me

홈 › Papers

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

2026-01-20 · Qian Chen, Jinlan Fu, Changsong Li, Min Zhang, See-Kiong Ng, Xipeng Qiu arxiv

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding. To bridge this gap, we introduce FutureOmni, the first benchmark designed to evaluate omni-modal future forecasting from audio-visual environments. The evaluated models are required to perform cross-modal causal and temporal reasoning, as well as effectively leverage internal knowledge to predict future events. FutureOmni is constructed via a scalable LLM-assisted, human-in-the-loop pipeline and contains 919 videos and 1,034 multiple-choice QA pairs across 8 primary domains. Evaluations on 13 omni-modal and 7 video-only models show that current systems struggle with audio-visual future prediction, particularly in speech-heavy scenarios, with the best accuracy of 64.8% achieved by Gemini 3 Flash. To mitigate this limitation, we curate a 7K-sample instruction-tuning dataset and propose an Omni-Modal Future Forecasting (OFF) training strategy. Evaluations on FutureOmni and popular audio-visual and video-only benchmarks demonstrate that OFF enhances future forecasting and generalization. We publicly release all code (https://github.com/OpenMOSS/FutureOmni) and datasets (https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni).

📄 PDF Abstract BibTeX arXiv:2601.13836

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

2025-06-26 · Yiman Zhang, Ziheng Luo, Qiangyu Yan, wei he 외

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has severa…

DiversityMultiple-choice

OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

2024-10-16 · Lichang Chen, Hexiang Hu, Mingda Zhang, YiWen Chen 외

We introduce OmnixR, an evaluation suite designed to benchmark SoTA Omni-modality Language Models, such as GPT-4o and Gemini. Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents…

OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales

2025-10-20 · Tung Nguyen, Tuan Pham, Troy Arcomano, Veerabhadra Kotamarthi 외 arxiv

Accurate weather forecasting across time scales is critical for anticipating and mitigating the impacts of climate change. Recent data-driven methods based on deep learning have achieved significant success in the medium…

Weather Forecasting

OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

2026-03-25 · Seunghee Kim, Bumkyu Park, Kyudan Jung, Joosung Lee 외 arxiv

Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we introduce OmniACBench, a benchmark for …

OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

2025-03-29 · CVPR 2025 1 · Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu 외

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despi…

Streaming video understandingVideo Understanding