paper-with-me

홈 › Papers

Frequency Domain Transformer Networks for Video Prediction

2019-03-01 · Hafez Farazi, Sven Behnke

The task of video prediction is forecasting the next frames given some previous frames. Despite much recent progress, this task is still challenging mainly due to high nonlinearity in the spatial domain. To address this issue, we propose a novel architecture, Frequency Domain Transformer Network (FDTN), which is an end-to-end learnable model that estimates and uses the transformations of the signal in the frequency domain. Experimental evaluations show that this approach can outperform some widely used video prediction methods like Video Ladder Network (VLN) and Predictive Gated Pyramids (PGP).

📄 PDF Abstract BibTeX arXiv:1903.00271

Code (1)

AIS-Bonn/FreqNet 공식 구현

Tasks

PredictionVideo Prediction

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Motion Segmentation using Frequency Domain Transformer Networks

2020-04-18 · Hafez Farazi, Sven Behnke

Self-supervised prediction is a powerful mechanism to learn representations that capture the underlying structure of the data. Despite recent progress, the self-supervised video prediction task is still challenging. One …

Motion SegmentationPredictionVideo Prediction

Frequency-Aware Spatiotemporal Transformers for Video Inpainting Detection

2021-01-01 · ICCV 2021 10 · Bingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 외

In this paper, we propose a frequency-aware spatiotemporal transformers for deep In this paper, we propose a Frequency-Aware Spatiotemporal Transformer (FAST) for video inpainting detection, which aims to simultaneou…

DecoderVideo Inpainting

Rethinking Urban Mobility Prediction: A Super-Multivariate Time Series Forecasting Approach

2023-12-04 · Jinguo Cheng, Ke Li, Yuxuan Liang, Lijun Sun 외

Long-term urban mobility predictions play a crucial role in the effective management of urban facilities and services. Conventionally, urban mobility data has been structured as spatiotemporal videos, treating longitude …

Multivariate Time Series ForecastingTime SeriesTime Series ForecastingVideo Prediction

Learning Spatiotemporal Frequency-Transformer for Low-Quality Video Super-Resolution

2022-12-27 · Zhongwei Qiu, Huan Yang, Jianlong Fu, Daochang Liu 외

Video Super-Resolution (VSR) aims to restore high-resolution (HR) videos from low-resolution (LR) videos. Existing VSR techniques usually recover HR frames by extracting pertinent textures from nearby frames with known d…

Super-ResolutionVideo EnhancementVideo Super-Resolution

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

2021-10-06 · Hafez Farazi, Jan Nogga, and Sven Behnke

The ultimate goal of video prediction is not forecasting future pixel-values given some previous frames. Rather, the end goal of video prediction is to discover valuable internal representations from the vast amount of a…

Decision MakingPredictionSemantic CompositionVideo Prediction