paper-with-me

홈 › Papers

Joint Reference Frame Synthesis and Post Filter Enhancement for Versatile Video Coding

2024-04-28 · Weijie Bao, Yuantong Zhang, Jianghao Jia, Zhenzhong Chen, Shan Liu

This paper presents the joint reference frame synthesis (RFS) and post-processing filter enhancement (PFE) for Versatile Video Coding (VVC), aiming to explore the combination of different neural network-based video coding (NNVC) tools to better utilize the hierarchical bi-directional coding structure of VVC. Both RFS and PFE utilize the Space-Time Enhancement Network (STENet), which receives two input frames with artifacts and produces two enhanced frames with suppressed artifacts, along with an intermediate synthesized frame. STENet comprises two pipelines, the synthesis pipeline and the enhancement pipeline, tailored for different purposes. During RFS, two reconstructed frames are sent into STENet's synthesis pipeline to synthesize a virtual reference frame, similar to the current to-be-coded frame. The synthesized frame serves as an additional reference frame inserted into the reference picture list (RPL). During PFE, two reconstructed frames are fed into STENet's enhancement pipeline to alleviate their artifacts and distortions, resulting in enhanced frames with reduced artifacts and distortions. To reduce inference complexity, we propose joint inference of RFS and PFE (JISE), achieved through a single execution of STENet. Integrated into the VVC reference software VTM-15.0, RFS, PFE, and JISE are coordinated within a novel Space-Time Enhancement Window (STEW) under Random Access (RA) configuration. The proposed method could achieve -7.34%/-17.21%/-16.65% PSNR-based BD-rate on average for three components under RA configuration.

📄 PDF Abstract BibTeX arXiv:2404.18058

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Recurrent Neural Network Postfilters for Statistical Parametric Speech Synthesis

2016-01-26 · Prasanna Kumar Muthukumar, Alan W. black

In the last two years, there have been numerous papers that have looked into using Deep Neural Networks to replace the acoustic model in traditional statistical parametric speech synthesis. However, far less attention ha…

General ClassificationregressionSpeech Synthesis

Latent CLAP Loss for Better Foley Sound Synthesis

2024-03-18 · Tornike Karchkhadze, Hassan Salami Kavaki, Mohammad Rasool Izadi, Bryce Irvin 외

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, su…

FAD

A neural network based post-filter for speech-driven head motion synthesis

2019-07-24 · JinHong Lu, Hiroshi Shimodaira

Despite the fact that neural networks are widely used for speech-driven head motion synthesis, it is well-known that the output of neural networks is noisy or discontinuous due to the limited capability of deep neural ne…

Motion Synthesis

Competitive Learning for Achieving Content-specific Filters in Video Coding for Machines

2024-06-18 · Honglei Zhang, Jukka I. Ahonen, Nam Le, Ruiying Yang 외

This paper investigates the efficacy of jointly optimizing content-specific post-processing filters to adapt a human oriented video/image codec into a codec suitable for machine vision tasks. By observing that artifacts …

Instance Segmentationobject-detectionObject DetectionSemantic Segmentation

WISE: Whitebox Image Stylization by Example-based Learning

2022-07-29 · Winfried Lötzsch, Max Reimann, Martin Büssemeyer, Amir Semmo 외

Image-based artistic rendering can synthesize a variety of expressive styles using algorithmic image filtering. In contrast to deep learning-based methods, these heuristics-based filtering techniques can operate on high-…

Image StylizationImage-to-Image TranslationParameter PredictionStyle Transfer