paper-with-me

홈 › Papers

NeMo: Needle in a Montage for Video-Language Understanding

2025-09-29 · Zi-Yuan Hu, Shuo Liang, Duo Zheng, Yanyang Li, Yeyao Tao, Shijia Huang, Wei Feng, Jia Qin, Jianguang Yu, Jing Huang, Meng Fang, Yin Li, Liwei Wang arxiv

Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for complex temporal reasoning in video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), designed to assess VideoLLMs' critical reasoning capabilities, including long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench.

📄 PDF Abstract BibTeX arXiv:2509.24563

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

Two Causally Related Needles in a Video Haystack

2025-05-26 · Miaoyu Li, Qin Chao, Boyang Li

Evaluating the video understanding capabilities of Video-Language Models (VLMs) remains a significant challenge. We propose a long-context video understanding benchmark, Causal2Needles, that assesses two crucial abilitie…

Video UnderstandingVisual Grounding

NVIDIA Nemotron Nano V2 VL

2025-11-06 · NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko 외 arxiv

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers…

Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

2024-06-13 · Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 외

Video understanding is a crucial next step for multimodal large language models (MLLMs). Various benchmarks are introduced for better evaluating the MLLMs. Nevertheless, current video benchmarks are still inefficient for…

BenchmarkingQuestion AnsweringVideo GenerationVideo Understanding+1

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

2026-04-27 · NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko 외 arxiv

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy im…

Autoregressive Modeling of Film with Applications in Video Montage

2026-07-16 · Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar, Bryan Rusell 외 arxiv

This work introduces FilmGPT, an autoregressive transformer designed to address the challenge of video montage--turning a collection of raw, "unwatchable" footage into coherent cinematic sequences. Inspired by language l…