paper-with-me

Papers

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

2026-08-30 · Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang, Dimitris Samaras, Vicky Kalogeiton arxiv

Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal large language models (MLLMs) are evaluated almost exclusively on understanding what happens rather than why it is presented that way. We introduce TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films (85 hours) with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent. Through controlled modality ablations, TAKE 85 enables systematic evaluation of multimodal reasoning. Experiments on state-of-the-art MLLMs reveal a substantial gap between perceptual recognition and intentional understanding: while models accurately describe events and narratives, they consistently fail to infer the communicative role of filmmaking decisions. Our results establish directorial intent as a previously overlooked dimension of multimodal understanding: even the strongest model reaches only 58 out of 100, and our ablations show that no input modality is sufficient on its own. All code, Q&As, and models are publicly available from https://github.com/KaiShinozakiConefrey/Take-85

📄 PDF Abstract BibTeX arXiv:2608.30068

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

On the Role of Noise in AudioVisual Integration: Evidence from Artificial Neural Networks that Exhibit the McGurk Effect

2024-11-08 · Lukas Grasse, Matthew S. Tata

Humans are able to fuse information from both auditory and visual modalities to help with understanding speech. This is frequently demonstrated through an phenomenon known as the McGurk Effect, during which a listener is…

Matching visual induction effects on screens of different size

2020-05-06 · Trevor D. Canham, Javier Vazquez-Corral, Elise Mathieu, Marcelo Bertalmío

In the film industry, the same movie is expected to be watched on displays of vastly different sizes, from cinema screens to mobile phones. But visual induction, the perceptual phenomenon by which the appearance of a sce…

Development and Evaluation of Video Recordings for the OLSA Matrix Sentence Test

2019-12-10 · Gerard Llorach, Frederike Kirschner, Giso Grimm, Melanie A. Zokoll 외

One of the established multi-lingual methods for testing speech intelligibility is the matrix sentence test (MST). Most versions of this test are designed with audio-only stimuli. Nevertheless, visual cues play an import…

Sentence

Modality Dropout for Improved Performance-driven Talking Faces

2020-05-27 · Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe 외

We describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-spee…

OWL (Observe, Watch, Listen): Audiovisual Temporal Context for Localizing Actions in Egocentric Videos

2022-02-10 · Merey Ramazanova, Victor Escorcia, Fabian Caba Heilbron, Chen Zhao 외

Egocentric videos capture sequences of human activities from a first-person perspective and can provide rich multimodal signals. However, most current localization methods use third-person videos and only incorporate vis…

Action LocalizationTemporal Action LocalizationTemporal Localization