paper-with-me

홈 › Papers

Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation

2025-09-06 · Xiaomeng Zhu, Changwei Wang, Haozhe Wang, Xinyu Liu, Fangzhen Lin arxiv

A scene graph is a structured representation of objects and their spatio-temporal relationships in dynamic scenes. Scene Graph Anticipation (SGA) involves predicting future scene graphs from video clips, enabling applications in intelligent surveillance and human-machine collaboration. While recent SGA approaches excel at leveraging visual evidence, long-horizon forecasting fundamentally depends on semantic priors and commonsense temporal regularities that are challenging to extract purely from visual features. To explicitly model these semantic dynamics, we propose Linguistic Scene Graph Anticipation (LSGA), a linguistic formulation of SGA that performs temporal relational reasoning over sequences of textualized scene graphs, with visual scene-graph detection handled by a modular front-end when operating on video. Building on this formulation, we introduce Object-Oriented Two-Stage Method (OOTSM), a language-based framework that anticipates object-set dynamics and forecasts object-centric relation trajectories with temporal consistency regularization, and we evaluate it on a dedicated benchmark constructed from Action Genome annotations. Extensive experiments show that compact fine-tuned language models with up to 3B parameters consistently outperform strong zero- and one-shot API baselines, including GPT-4o, GPT-4o-mini, and DeepSeek-V3, under matched textual inputs and context windows. When coupled with off-the-shelf visual scene-graph generators, the resulting multimodal system achieves substantial improvements on video-based SGA, boosting long-horizon mR@50 by up to 21.9\% over strong visual SGA baselines.

📄 PDF Abstract BibTeX arXiv:2509.05661

Code (0)

등록된 구현이 없습니다.

Tasks

Relational Reasoning

Similar Papers 제목 키워드 기반

Task-Oriented Human Grasp Synthesis via Context- and Task-Aware Diffusers

2025-07-15 · An-Lun Liu, Yu-Wei Chao, Yi-Ting Chen

In this paper, we study task-oriented human grasp synthesis, a new grasp synthesis task that demands both task and context awareness. At the core of our method is the task-aware contact maps. Unlike traditional contact m…

Oriented Objects as pairs of Middle Lines

2019-12-23 · Hao-Ran Wei, Yue Zhang, Zhonghan Chang, Hao Li 외

The detection of oriented objects is frequently appeared in the field of natural scene text detection as well as object detection in aerial images. Traditional detectors for oriented objects are common to rotate anchors …

object-detectionObject DetectionObject Detection In Aerial ImagesOne-stage Anchor-free Oriented Object Detection+3

LOGOS: Language-guided Oriented Object Detection in Aerial Scenes

2026-07-09 · Trong-Thuan Nguyen, Minh-Triet Tran arxiv

Object detection in geospatial scenes, such as satellite and aerial imagery, poses significant challenges due to the varying orientations and densities of objects, as well as the complex backgrounds inherent to remote se…

Object Detection

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models

2026-05-13 · Vladislav Makarov, Mark Gizetdinov, Dmitry Yudin arxiv

Scene graph generation provides a compact structured representation for visual perception, but accurate and fast graph prediction from images and videos remains challenging. Recent VLM-based methods can generate scene gr…

Video scene graph generationReinforcement Learning

ReSpace: Text-Driven 3D Scene Synthesis and Editing with Preference Alignment

2025-06-03 · Martin JJ. Bucher, Iro Armeni

Scene synthesis and editing has emerged as a promising direction in computer graphics. Current trained approaches for 3D indoor scenes either oversimplify object semantics through one-hot class encodings (e.g., 'chair' o…

Indoor Scene SynthesisObjectSpatial Reasoning