paper-with-me

홈 › Papers

Complex Sequential Understanding through the Awareness of Spatial and Temporal Concepts

2020-05-30 · Bo Pang, Kaiwen Zha, Hanwen Cao, Jiajun Tang, Minghui Yu, Cewu Lu

Understanding sequential information is a fundamental task for artificial intelligence. Current neural networks attempt to learn spatial and temporal information as a whole, limited their abilities to represent large scale spatial representations over long-range sequences. Here, we introduce a new modeling strategy called Semi-Coupled Structure (SCS), which consists of deep neural networks that decouple the complex spatial and temporal concepts learning. Semi-Coupled Structure can learn to implicitly separate input information into independent parts and process these parts respectively. Experiments demonstrate that a Semi-Coupled Structure can successfully annotate the outline of an object in images sequentially and perform video action recognition. For sequence-to-sequence problems, a Semi-Coupled Structure can predict future meteorological radar echo images based on observed images. Taken together, our results demonstrate that a Semi-Coupled Structure has the capacity to improve the performance of LSTM-like models on large scale sequential tasks.

📄 PDF Abstract BibTeX arXiv:2006.00212

Code (1)

BoPang1996/Semi-Coupled-Structure-for-visual-sequental-tasks 공식 구현 pytorch

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

2026-07-14 · Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan 외 arxiv

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also ma…

RegionGPT: Towards Region Understanding Vision Language Model

2024-03-04 · CVPR 2024 1 · Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon 외

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limite…

Language ModelingLanguage Modellingmodel

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

2025-03-25 · Kexian Tang, Junyao Gao, Yanhong Zeng, Haodong Duan 외

Multi-step spatial reasoning entails understanding and reasoning about spatial relationships across multiple sequential steps, which is crucial for tackling complex real-world applications, such as robotic manipulation, …

Autonomous NavigationQuestion AnsweringSpatial ReasoningVisual Question Answering+1

Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation

2026-02-12 · Wei Chen, Yancheng Long, Mingqiao Liu, Haojie Ding 외 arxiv

While diffusion models have shown exceptional capabilities in aesthetic image synthesis, they often struggle with complex spatial understanding and reasoning. Existing approaches resort to Multimodal Large Language Model…

Spatial ReasoningImage GenerationImage Editing

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness

2026-04-29 · Haiyi Qiu, Kaihang Pan, Jiacheng Li, Juncheng Li 외 arxiv

Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image generation. However, these models remain fundamentally limited in sp…

Text-to-Image GenerationImage Editing