paper-with-me

Papers

Fine-grained spatial-temporal perception for gas leak segmentation

2025-05-01 · Xinlong Zhao, Shan Du

Gas leaks pose significant risks to human health and the environment. Despite long-standing concerns, there are limited methods that can efficiently and accurately detect and segment leaks due to their concealed appearance and random shapes. In this paper, we propose a Fine-grained Spatial-Temporal Perception (FGSTP) algorithm for gas leak segmentation. FGSTP captures critical motion clues across frames and integrates them with refined object features in an end-to-end network. Specifically, we first construct a correlation volume to capture motion information between consecutive frames. Then, the fine-grained perception progressively refines the object-level features using previous outputs. Finally, a decoder is employed to optimize boundary segmentation. Because there is no highly precise labeled dataset for gas leak segmentation, we manually label a gas leak video dataset, GasVid. Experimental results on GasVid demonstrate that our model excels in segmenting non-rigid objects such as gas leaks, generating the most accurate mask compared to other state-of-the-art (SOTA) models.

📄 PDF Abstract BibTeX arXiv:2505.00295

Code (1)

GeekEagle/FGSTP 공식 구현 pytorch

Tasks

DecoderSegmentation

Similar Papers 제목 키워드 기반

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

2026-06-26 · Yankai Yang, Yancheng Long, Bin Wen, Fan Yang 외 arxiv

Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and …

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

2026-05-31 · Xin Dong, Wenjia Geng, Wenfeng Deng, Yansong Tang arxiv

Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat the…

Representation LearningHighlight DetectionMoment RetrievalVideo Grounding

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

2025-09-15 · Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju 외 arxiv

Recent advancements in large video models (LVMs) have significantly enhance video understanding. However, these models continue to suffer from hallucinations, producing content that conflicts with input videos. To addres…

Fine-grained Context and Multi-modal Alignment for Freehand 3D Ultrasound Reconstruction

2024-07-05 · Zhongnuo Yan, Xin Yang, Mingyuan Luo, Jiongquan Chen 외

Fine-grained spatio-temporal learning is crucial for freehand 3D ultrasound reconstruction. Previous works mainly resorted to the coarse-grained spatial features and the separated temporal dependency learning and struggl…

Management

Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

2024-05-07 · Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang 외

Existing research of video understanding still struggles to achieve in-depth comprehension and reasoning in complex videos, primarily due to the under-exploration of two key bottlenecks: fine-grained spatial-temporal per…

Large Language ModelMultimodal Large Language ModelVideo GroundingVideo Understanding