paper-with-me

홈 › Papers

From Benedict Cumberbatch to Sherlock Holmes: Character Identification in TV series without a Script

2018-01-31 · Arsha Nagrani, Andrew Zisserman

The goal of this paper is the automatic identification of characters in TV and feature film material. In contrast to standard approaches to this task, which rely on the weak supervision afforded by transcripts and subtitles, we propose a new method requiring only a cast list. This list is used to obtain images of actors from freely available sources on the web, providing a form of partial supervision for this task. In using images of actors to recognize characters, we make the following three contributions: (i) We demonstrate that an automated semi-supervised learning approach is able to adapt from the actor's face to the character's face, including the face context of the hair; (ii) By building voice models for every character, we provide a bridge between frontal faces (for which there is plenty of actor-level supervision) and profile (for which there is very little or none); and (iii) by combining face context and speaker identification, we are able to identify characters with partially occluded faces and extreme facial poses. Results are presented on the TV series 'Sherlock' and the feature film 'Casablanca'. We achieve the state-of-the-art on the Casablanca benchmark, surpassing previous methods that have used the stronger supervision available from transcripts.

📄 PDF Abstract BibTeX arXiv:1801.10442

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Identification

Similar Papers 제목 키워드 기반

Report on the First Knowledge Graph Reasoning Challenge 2018 -- Toward the eXplainable AI System

2019-08-22 · Takahiro Kawamura, Shusaku Egami, Koutarou Tamura, Yasunori Hokazono 외

A new challenge for knowledge graph reasoning started in 2018. Deep learning has promoted the application of artificial intelligence (AI) techniques to a wide variety of social problems. Accordingly, being able to explai…

graph construction

Illustrating a neural model of logic computations: The case of Sherlock Holmes' old maxim

2012-10-28 · Eduardo Mizraji

Natural languages can express some logical propositions that humans are able to understand. We illustrate this fact with a famous text that Conan Doyle attributed to Holmes: 'It is an old maxim of mine that when you have…

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

2025-05-27 · Junhao Cheng, Yuying Ge, Teng Wang, Yixiao Ge 외

Recent advances in CoT reasoning and RL post-training have been reported to enhance video reasoning capabilities of MLLMs. This progress naturally raises a question: can these models perform complex video reasoning in a …

Multimodal Reasoning

Characterizing the Investigative Methods of Fictional Detectives with Large Language Models

2025-05-12 · Edirlei Soares de Lima, Marco A. Casanova, Bruno Feijó, Antonio L. Furtado

Detective fiction, a genre defined by its complex narrative structures and character-driven storytelling, presents unique challenges for computational narratology, a research field focused on integrating literary theory …

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework

2025-05-25 · Feiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng 외

Our research reveals a new privacy risk associated with the vision-language model (VLM) agentic framework: the ability to infer sensitive attributes (e.g., age and health information) and even abstract ones (e.g., person…

AttributeLanguage ModelingLanguage ModellingVisual Reasoning