paper-with-me

홈 › Papers

Deep Multimodal Feature Encoding for Video Ordering

2020-04-05 · Vivek Sharma, Makarand Tapaswi, Rainer Stiefelhagen

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that encodes all these modalities. Our model parameters are learned through a proxy task of inferring the temporal ordering of a set of unordered videos in a timeline. To this end, we create a new multimodal dataset for temporal ordering that consists of approximately 30K scenes (2-6 clips per scene) based on the "Large Scale Movie Description Challenge". We analyze and evaluate the individual and joint modalities on three challenging tasks: (i) inferring the temporal ordering of a set of videos; and (ii) action recognition. We demonstrate empirically that multimodal representations are indeed complementary, and can play a key role in improving the performance of many applications.

📄 PDF Abstract BibTeX arXiv:2004.02205

Code (1)

vivoutlaw/tcbp 공식 구현 pytorch

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020

2020-06-23 · Tosho Hirasawa, Zhishen Yang, Mamoru Komachi, Naoaki Okazaki

Video-guided machine translation as one of multimodal neural machine translation tasks targeting on generating high-quality text translation by tangibly engaging both video and text. In this work, we presented our video-…

Machine TranslationTranslationVideo-Guided Machine Translation

Multimodal Semantic Attention Network for Video Captioning

2019-05-08 · Liang Sun, Bing Li, Chunfeng Yuan, Zheng-Jun Zha 외

Inspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder framework incorporating multimodal semant…

AttributeDecoderGeneral ClassificationMulti-Label Classification+2

Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

2026-07-23 · Liu Liu, Freya Huying Tan, Fábio Duarte arxiv

We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented acro…

Binary Classification

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

2026-03-30 · Milton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen 외 arxiv

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high pr…

CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models

2026-01-08 · Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto 외 arxiv

Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitigation strategies, such as textual filteri…

Action RecognitionVideo Generation