paper-with-me

Papers

Structure-Preserving Patch Decoding for Efficient Neural Video Representation

2025-06-15 · Taiga Hayami, Kakeru Koizumi, Hiroshi Watanabe

Implicit neural representations (INRs) are the subject of extensive research, particularly in their application to modeling complex signals by mapping spatial and temporal coordinates to corresponding values. When handling videos, mapping compact inputs to entire frames or spatially partitioned patch images is an effective approach. This strategy better preserves spatial relationships, reduces computational overhead, and improves reconstruction quality compared to coordinate-based mapping. However, predicting entire frames often limits the reconstruction of high-frequency visual details. Additionally, conventional patch-based approaches based on uniform spatial partitioning tend to introduce boundary discontinuities that degrade spatial coherence. We propose a neural video representation method based on Structure-Preserving Patches (SPPs) to address such limitations. Our method separates each video frame into patch images of spatially aligned frames through a deterministic pixel-based splitting similar to PixelUnshuffle. This operation preserves the global spatial structure while allowing patch-level decoding. We train the decoder to reconstruct these structured patches, enabling a global-to-local decoding strategy that captures the global layout first and refines local details. This effectively reduces boundary artifacts and mitigates distortions from naive upsampling. Experiments on standard video datasets demonstrate that our method achieves higher reconstruction quality and better compression performance than existing INR-based baselines.

📄 PDF Abstract BibTeX arXiv:2506.12896

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PS-NeRV: Patch-wise Stylized Neural Representations for Videos

2022-08-07 · Yunpeng Bai, Chao Dong, Cairong Wang

We study how to represent a video with implicit neural representations (INRs). Classical INRs methods generally utilize MLPs to map input coordinates to output pixels. While some recent works have tried to directly recon…

Video CompressionVideo InpaintingVideo Reconstruction

NIRVANA: Neural Implicit Representations of Videos with Adaptive Networks and Autoregressive Patch-wise Modeling

2022-12-30 · CVPR 2023 1 · Shishira R Maiya, Sharath Girish, Max Ehrlich, Hanyu Wang 외

Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exploit the temporal redundancy in videos, …

QuantizationVideo Compression

Rethinking Patch Dependence for Masked Autoencoders

2024-01-25 · Letian Fu, Long Lian, Renhao Wang, Baifeng Shi 외

In this work, we re-examine inter-patch dependencies in the decoding mechanism of masked autoencoders (MAE). We decompose this decoding mechanism for masked patch reconstruction in MAE into self-attention and cross-atten…

DecoderInstance SegmentationRepresentation LearningSemantic Segmentation

DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers

2026-02-19 · Dahye Kim, Deepti Ghadiyaram, Raghudeep Gadde arxiv

Diffusion Transformers (DiTs) have achieved state-of-the-art performance in image and video generation, but their success comes at the cost of heavy computation. This inefficiency is largely due to the fixed tokenization…

Video Generation

Real-Time Anomalous Behavior Detection and Localization in Crowded Scenes

2015-11-21 · Mohammad Sabokrou, Mahmood Fathy, Mojtaba Hosseini

In this paper, we propose an accurate and real-time anomaly detection and localization in crowded scenes, and two descriptors for representing anomalous behavior in video are proposed. We consider a video as being a set …

Anomaly Detection