Character Grounding and Re-Identification in Story of Videos and Text Descriptions
We address character grounding and re-identification in multiple story-based videos like movies and associated text descriptions. In order to solve these related tasks in a mutually rewarding way, we propose a model named Character in Story Identification Network (CiSIN). Our method builds two semantically informative representations via joint training of multiple objectives for character grounding, video/text re-identification and gender prediction: Visual Track Embedding from videos and Textual Character Embedding from text context. These two representations are learned to retain rich semantic multimodal information that enables even simple MLPs to achieve the state-of-the-art performance on the target tasks. More specifically, our CiSIN model achieves the best performance in the Fill-in the Characters task of LSMDC 2019 challenges. Moreover, it outperforms previous state-of-the-art models in M-VAD Names dataset as a benchmark of multimodal character grounding and re-identification.
Code (0)
등록된 구현이 없습니다.
Tasks
Gender PredictionSimilar Papers 제목 키워드 기반
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations. These issues can be addressed through groundin…
Face RecognitionObjectobject-detectionObject Detection+3StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce Story…
Visual StorytellingVisual GroundingEntity Re-identification in Visual Storytelling via Contrastive Reinforcement Learning
Visual storytelling systems, particularly large vision-language models, struggle to maintain character and object identity across frames, often failing to recognize when entities in different images represent the same in…
Reinforcement LearningVisual StorytellingStoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…
Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo DescriptionDetecting and Grounding Important Characters in Visual Stories
Characters are essential to the plot of any story. Establishing the characters before writing a story can improve the clarity of the plot and the overall flow of the narrative. However, previous work on visual storytelli…
Visual Storytelling