paper-with-me

홈 › Papers

Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models

2026-09-15 · Rwiddhi Chakraborty, Yinong, Wang, Cheng Zhang, Fan Bai, Zhuoran You, Michael Kampffmeyer, Yong Jae Lee, Fernando De la Torre, Robert Jenssen arxiv

Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also been demonstrated on a wide range of benchmarks, with an important caveat - the dominant approach in these benchmarks evaluates multiple choice reasoning via text options. This is a natural way to test text-based reasoning in these models, and has led to significant insights regarding model behavior in the community. In this work, we ask a different question - what happens when the evaluation modality is visual, rather than text? We introduce three new vision-centric evaluation benchmarks in temporal frame retrieval, video future prediction, and causal memory distortion, all designed around evaluating visual understanding capabilities in large video models. Our approach complements the existing approaches to evaluate video understanding in frontier models. We show that current frontier models exhibit significant weakness when attempting to reason through visual queries, rather than text. We conclude with an extended analysis section that provides pointers for future improvements in visual understanding for large video models.

📄 PDF Abstract BibTeX arXiv:2609.17112

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Integrating Parametric and Non-Parametric Models For Scene Labeling

2015-06-01 · CVPR 2015 6 · Bing Shuai, Gang Wang, Zhen Zuo, Bing Wang 외

We adopt Convolutional Neural Networks (CNN) as our parametric model to learn discriminative features and classifiers for local patch classification. As visually similar pixels are indistinguishable from local context, w…

General ClassificationMetric LearningScene Labeling

SLAM-Former: Putting SLAM into One Transformer

2025-09-21 · Yijun Yuan, Zhuoguang Chen, Kenan Li, Weibang Wang 외 arxiv

We present SLAM-Former, a novel neural approach that integrates full SLAM capabilities into a single transformer. Similar to traditional SLAM systems, SLAM-Former comprises both a frontend and a backend that operate in t…

Deductive Association Networks

2021-11-02 · Seokjun Kim, Jaeeun Jang, Hyeoncheol Kim

we introduce deductive association networks(DANs), a network that performs deductive reasoning. To have high-dimensional thinking, combining various axioms and putting the results back into another axiom is necessary to …

Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution

2021-01-24 · Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang 외

Visual information extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and intelligent education. Most existing work…

3D Feature Matchingdocument understandingText DetectionText Spotting

Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach

2024-01-28 · Shaofeng Zhang, Jinfa Huang, Qiang Zhou, Zhibin Wang 외

Image outpainting aims to generate the content of an input sub-image beyond its original boundaries. It is an important task in content generation yet remains an open problem for generative models. This paper pushes the …

Image Outpainting