paper-with-me

Papers

Exploiting Temporal Coherence for Multi-modal Video Categorization

2020-02-07 · Palash Goyal, Saurabh Sahu, Shalini Ghosh, Chul Lee

Multimodal ML models can process data in multiple modalities (e.g., video, images, audio, text) and are useful for video content analysis in a variety of problems (e.g., object detection, scene understanding). In this paper, we focus on the problem of video categorization by using a multimodal approach. We have developed a novel temporal coherence-based regularization approach, which applies to different types of models (e.g., RNN, NetVLAD, Transformer). We demonstrate through experiments how our proposed multimodal video categorization models with temporal coherence out-perform strong state-of-the-art baseline models.

📄 PDF Abstract BibTeX arXiv:2002.03844

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionScene Understanding

Similar Papers 제목 키워드 기반

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

2025-03-12 · Xiaowei Bi, Zheyuan Xu

Long Video Question Answering (LVQA) is challenging due to the need for temporal reasoning and large-scale multimodal data processing. Existing methods struggle with retrieving cross-modal information from long videos, e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Information RetrievalQuestion Answering+8

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

2026-04-11 · Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan arxiv

Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantica…

Video Summarization

EmoCo: Visual Analysis of Emotion Coherence in Presentation Videos

2019-07-29 · Haipeng Zeng, Xingbo Wang, Aoyu Wu, Yong Wang 외

Emotions play a key role in human communication and public presentations. Human emotions are usually expressed through multiple modalities. Therefore, exploring multimodal emotions and their coherence is of great value f…

ClusteringSentence

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

2024-12-16 · Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar 외

Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we p…

InformativenessLarge Language ModelText GenerationText-to-Video Generation+2

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

2026-02-10 · Jialun Liu, Tian Li, Xiao Cao, Yukuo Ma 외 arxiv

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, l…

Text-to-Video Generation