paper-with-me

홈 › Papers

From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics

2026-03-24 · Paolo Cupini, Francesco Pierri arxiv

Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, domain-specific editorial patterns, and strict operational constraints. While multimodal large language models (MLLMs) have demonstrated strong general-purpose video understanding capabilities, their comparative effectiveness across pipeline architectures and input configurations in broadcast-specific settings remains empirically undercharacterized. This paper presents a systematic evaluation of multimodal annotation pipelines applied to broadcast television news in the Italian setting. We construct a domain-specific benchmark of clips labeled across four semantic dimensions: visual environment classification, topic classification, sensitive content detection, and named entity recognition. Two different pipeline architectures are evaluated across nine frontier models, including Gemini 3.0 Pro, LLaMA 4 Maverick, Qwen-VL variants, and Gemma 3, under progressively enriched input strategies combining visual signals, automatic speech recognition, speaker diarization, and metadata. Experimental results demonstrate that gains from video input are strongly model-dependent: larger models effectively leverage temporal continuity, while smaller models show performance degradation under extended multimodal context, likely due to token overload. Beyond benchmarking, the selected pipeline is deployed on 14 full broadcast episodes, with minute-level annotations integrated with normalized audience measurement data provided by an Italian media company. This integration enables correlational analysis of topic-level audience sensitivity and generational engagement divergence, demonstrating the operational viability of the proposed framework for content-based audience analytics.

📄 PDF Abstract BibTeX arXiv:2603.26772

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech Recognition

Similar Papers 제목 키워드 기반

Deep Live Video Ad Placement on the 5G Edge

2022-09-03 · Mohammad Hosseini

The video broadcasting industry has been growing significantly in the recent years, specially on delivering personalized contents to the end users. While video broadcasting has continued to grow beyond TV, video advertin…

Marketing

Game-MUG: Multimodal Oriented Game Situation Understanding and Commentary Generation Dataset

2024-04-30 · Zhihao Zhang, Feiqi Cao, Yingbin Mo, Yiran Zhang 외

The dynamic nature of esports makes the situation relatively complicated for average viewers. Esports broadcasting involves game expert casters, but the caster-dependent game commentary is not enough to fully understand …

Time Series

From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework

2026-05-01 · Zihan Ding, Ziyuan Yang, Yi Zhang arxiv

Multimodal controversy detection (MCD) identifies controversial content in videos and their associated user comments, to support risk management for social video platforms.Prior research frames MCD as a static representa…

Representation Learning

MultiVENT: Multilingual Videos of Events with Aligned Natural Text

2023-07-06 · Kate Sanders, David Etter, Reno Kriz, Benjamin Van Durme

Everyday news coverage has shifted from traditional broadcasts towards a wide range of presentation formats such as first-hand, unedited video footage. Datasets that reflect the diverse array of multimodal, multilingual …

Information RetrievalRetrievalVideo Retrieval

MultiVENT: Multilingual Videos of Events and Aligned Natural Text

2023-09-26 · NeurIPS 2023 11

Everyday news coverage has shifted from traditional broadcasts towards a wide range of presentation formats such as first-hand, unedited video footage. Datasets that reflect the diverse array of multimodal, multilingual …