paper-with-me

Papers

Tarsier: Evolving Noise Injection in Super-Resolution GANs

2020-09-25 · Baptiste Roziere, Nathanal Carraz Rakotonirina, Vlad Hosu, Andry Rasoanaivo, Hanhe Lin, Camille Couprie, Olivier Teytaud

Super-resolution aims at increasing the resolution and level of detail within an image. The current state of the art in general single-image super-resolution is held by NESRGAN+, which injects a Gaussian noise after each residual layer at training time. In this paper, we harness evolutionary methods to improve NESRGAN+ by optimizing the noise injection at inference time. More precisely, we use Diagonal CMA to optimize the injected noise according to a novel criterion combining quality assessment and realism. Our results are validated by the PIRM perceptual score and a human study. Our method outperforms NESRGAN+ on several standard super-resolution datasets. More generally, our approach can be used to optimize any method based on noise injection.

📄 PDF Abstract BibTeX arXiv:2009.12177

Code (1)

ncarraz/ESRGANplus pytorch

Tasks

Image Super-ResolutionSuper-Resolution

Similar Papers 제목 키워드 기반

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

2025-01-14 · Liping Yuan, Jiawei Wang, Haomiao Sun, Yuchen Zhang 외

We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tars…

Embodied Question AnsweringHallucinationLanguage ModelingLanguage Modelling+5

Tarsier: Recipes for Training and Evaluating Large Video Description Models

2024-06-30 · arXiv 2024 7 · Jiawei Wang, Liping Yuan, Yuchen Zhang, Haomiao Sun

Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models designed to generate high-quality video desc…

Video CaptioningVideo DescriptionVideo Question AnsweringVideo Understanding+2

https://arxiv.org/abs/2407.00634

2024-07-02 · Jiawei Wang, Liping Yuan, Yuchen Zhang

Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models designed to generate high-quality video desc…

Video CaptioningVideo DescriptionVideo UnderstandingVisual Question Answering (VQA)

Real-World Super-Resolution via Kernel Estimation and Noise Injection

2020-06-19 · CVPRW 2020 6 · Xiaozhong Ji, Yun Cao, Ying Tai, Chengjie Wang 외

Recent state-of-the-art super-resolution methods have achieved impressive performance on ideal datasets regardless of blur and noise. However, these methods always fail in real-world image super-resolution, since most of…

Image Super-ResolutionSuper-ResolutionVideo Super-Resolution

When Large Multimodal Models Confront Evolving Knowledge:Challenges and Pathways

2025-05-30 · Kailin Jiang, Yuntao Du, Yukai Ding, Yuchen Ren 외

Large language/multimodal models (LLMs/LMMs) store extensive pre-trained knowledge but struggle to maintain consistency with real-world updates, making it difficult to avoid catastrophic forgetting while acquiring evolvi…

Continual LearningImage AugmentationInstruction Following