paper-with-me

Papers

Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution

2026-03-09 · Hua Liu, Yanbin Wei, Fei Xing, Tyler Derr, Haoyu Han, Yu Zhang arxiv

Dynamic graphs are common in real-world systems such as social media, recommender systems, and traffic networks. Existing dynamic graph models for link prediction often fall short in capturing the complexity of temporal evolution. They tend to overlook fine-grained variations in temporal interaction order, struggle with dependencies that span long time horizons, and offer limited capability to model pair-specific relational dynamics. To address these challenges, we propose \textbf{Graph2Video}, a video-inspired framework that views the temporal neighborhood of a target link as a sequence of "graph frames". By stacking temporally ordered subgraph frames into a "graph video", Graph2Video leverages the inductive biases of video foundation models to capture both fine-grained local variations and long-range temporal dynamics. It generates a link-level embedding that serves as a lightweight and plug-and-play link-centric memory unit. This embedding integrates seamlessly into existing dynamic graph encoders, effectively addressing the limitations of prior approaches. Extensive experiments on benchmark datasets show that Graph2Video outperforms state-of-the-art baselines on the link prediction task in most cases. The results highlight the potential of borrowing spatio-temporal modeling techniques from computer vision as a promising and effective approach for advancing dynamic graph learning.

📄 PDF Abstract BibTeX arXiv:2603.13360

Code (0)

등록된 구현이 없습니다.

Tasks

Link PredictionGraph Learning

Similar Papers 제목 키워드 기반

(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering

2022-02-18 · Anoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le Roux

Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are …

Question AnsweringSpatio-temporal Scene GraphsVideo Question Answering

Event-Enhanced Snapshot Compressive Videography at 10K FPS

2024-04-11 · Bo Zhang, Jinli Suo, Qionghai Dai

Video snapshot compressive imaging (SCI) encodes the target dynamic scene compactly into a snapshot and reconstructs its high-speed frame sequence afterward, greatly reducing the required data footprint and transmission …

Video Frame Interpolation

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

2025-05-21 · Fatemeh Ziaeetabar, Florentin Wörgötter

Foundation models have ushered in a new era for multimodal video understanding by enabling the extraction of rich spatiotemporal and semantic representations. In this work, we introduce a novel graph-based framework that…

Action RecognitionGraph AttentionVideo Understanding

Poet: Product-oriented Video Captioner for E-commerce

2020-08-16 · Shengyu Zhang, Ziqi Tan, Jin Yu, Zhou Zhao 외

In e-commerce, a growing number of user-generated videos are used for product promotion. How to generate video descriptions that narrate the user-preferred product characteristics depicted in the video is vital for succe…

Video Captioning

Fundus2Video: Cross-Modal Angiography Video Generation from Static Fundus Photography with Clinical Knowledge Guidance

2024-08-27 · Weiyi Zhang, Siyu Huang, Jiancheng Yang, Ruoyu Chen 외

Fundus Fluorescein Angiography (FFA) is a critical tool for assessing retinal vascular dynamics and aiding in the diagnosis of eye diseases. However, its invasive nature and less accessibility compared to Color Fundus (C…

Clinical KnowledgeLesion SegmentationVideo Generation