paper-with-me

홈 › Papers

A Unified Model for Video Understanding and Knowledge Embedding with Heterogeneous Knowledge Graph Dataset

2022-11-19 · Jiaxin Deng, Dong Shen, Haojie Pan, Xiangyu Wu, Ximan Liu, Gaofeng Meng, Fan Yang, Size Li, Ruiji Fu, Zhongyuan Wang

Video understanding is an important task in short video business platforms and it has a wide application in video recommendation and classification. Most of the existing video understanding works only focus on the information that appeared within the video content, including the video frames, audio and text. However, introducing common sense knowledge from the external Knowledge Graph (KG) dataset is essential for video understanding when referring to the content which is less relevant to the video. Owing to the lack of video knowledge graph dataset, the work which integrates video understanding and KG is rare. In this paper, we propose a heterogeneous dataset that contains the multi-modal video entity and fruitful common sense relations. This dataset also provides multiple novel video inference tasks like the Video-Relation-Tag (VRT) and Video-Relation-Video (VRV) tasks. Furthermore, based on this dataset, we propose an end-to-end model that jointly optimizes the video understanding objective with knowledge graph embedding, which can not only better inject factual knowledge into video understanding but also generate effective multi-modal entity embedding for KG. Comprehensive experiments indicate that combining video understanding embedding with factual knowledge benefits the content-based video retrieval performance. Moreover, it also helps the model generate better knowledge graph embedding which outperforms traditional KGE-based methods on VRT and VRV tasks with at least 42.36% and 17.73% improvement in HITS@10.

📄 PDF Abstract BibTeX arXiv:2211.10624

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningGraph EmbeddingKnowledge Graph EmbeddingRetrievalTAGVideo RetrievalVideo Understanding

Similar Papers 제목 키워드 기반

CogKGE: A Knowledge Graph Embedding Toolkit and Benchmark for Representing Multi-source and Heterogeneous Knowledge

2022-05-01 · ACL 2022 5 · Zhuoran Jin, Tianyi Men, Hongbang Yuan, Zhitao He 외

In this paper, we propose CogKGE, a knowledge graph embedding (KGE) toolkit, which aims to represent multi-source and heterogeneous knowledge. For multi-source knowledge, unlike existing methods that mainly focus on enti…

Graph EmbeddingKnowledge Graph Embedding

DreamWorld: Unified World Modeling in Video Generation

2026-02-28 · Boming Tan, Xiangdong Zhang, Ning Liao, Yuqing Zhang 외 arxiv

Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a sing…

Video Generation

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

2026-06-06 · Qi Liu, Gang Yue, Mingyu Yin, Lisai Zhang 외 arxiv

Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Building a unified framework that supports…

Video Generation

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

2025-12-08 · Jiehui Huang, Yuechen Zhang, Xu He, Yuan Gao 외 arxiv

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal i…

Zero-shot GeneralizationMulti-Task LearningVideo Generation

Bernini: Latent Semantic Planning for Video Diffusion

2026-05-21 · Bernini Team, Chenchen Liu, Junyi Chen, Lei Li 외 arxiv

Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong semantic grounding, while diffusion models sy…

Video Generation