paper-with-me

Papers

DIffSteISR: Harnessing Diffusion Prior for Superior Real-world Stereo Image Super-Resolution

2024-08-14 · Yuanbo Zhou, Xinlin Zhang, Wei Deng, Tao Wang, Tao Tan, Qinquan Gao, Tong Tong

We introduce DiffSteISR, a pioneering framework for reconstructing real-world stereo images. DiffSteISR utilizes the powerful prior knowledge embedded in pre-trained text-to-image model to efficiently recover the lost texture details in low-resolution stereo images. Specifically, DiffSteISR implements a time-aware stereo cross attention with temperature adapter (TASCATA) to guide the diffusion process, ensuring that the generated left and right views exhibit high texture consistency thereby reducing disparity error between the super-resolved images and the ground truth (GT) images. Additionally, a stereo omni attention control network (SOA ControlNet) is proposed to enhance the consistency of super-resolved images with GT images in the pixel, perceptual, and distribution space. Finally, DiffSteISR incorporates a stereo semantic extractor (SSE) to capture unique viewpoint soft semantic information and shared hard tag semantic information, thereby effectively improving the semantic accuracy and consistency of the generated left and right images. Extensive experimental results demonstrate that DiffSteISR accurately reconstructs natural and precise textures from low-resolution stereo images while maintaining a high consistency of semantic and texture between the left and right views.

📄 PDF Abstract BibTeX arXiv:2408.07516

Code (0)

등록된 구현이 없습니다.

Tasks

Image Super-ResolutionStereo Image Super-ResolutionSuper-ResolutionTAG

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Adapter 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CAD: Photorealistic 3D Generation via Adversarial Distillation

2023-12-11 · CVPR 2024 1 · Ziyu Wan, Despoina Paschalidou, IAn Huang, Hongyu Liu 외

The increased demand for 3D data in AR/VR, robotics and gaming applications, gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation…

3D GenerationDiversity

HIPPo: Harnessing Image-to-3D Priors for Model-free Zero-shot 6D Pose Estimation

2025-02-14 · Yibo Liu, Zhaodong Jiang, Binbin Xu, Guile Wu 외

This work focuses on model-free zero-shot 6D object pose estimation for robotics applications. While existing methods can estimate the precise 6D pose of objects, they heavily rely on curated CAD models or reference imag…

3D Reconstruction6D Pose Estimation6D Pose Estimation using RGBImage to 3D+2

Arbitrary-steps Image Super-resolution via Diffusion Inversion

2024-12-12 · CVPR 2025 1 · Zongsheng Yue, Kang Liao, Chen Change Loy

This study presents a new image super-resolution (SR) technique based on diffusion inversion, aiming at harnessing the rich image priors encapsulated in large pre-trained diffusion models to improve SR performance. We de…

Image Super-ResolutionSuper-Resolution

NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer

2024-05-24 · Meng You, Zhiyu Zhu, Hui Liu, Junhui Hou

By harnessing the potent generative capabilities of pre-trained large video diffusion models, we propose NVS-Solver, a new novel view synthesis (NVS) paradigm that operates \textit{without} the need for training. NVS-Sol…

Novel View Synthesis

ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion

2025-09-09 · Ao Li, Jinpeng Liu, Yixuan Zhu, Yansong Tang arxiv

Joint reconstruction of human-object interaction marks a significant milestone in comprehending the intricate interrelations between humans and their surrounding environment. Nevertheless, previous optimization methods o…