paper-with-me

Papers

Self-Refining Video Sampling

2026-01-26 · Sangwon Jang, Taekyung Ki, Jaehyeong Jo, Saining Xie, Jaehong Yoon, Sung Ju Hwang arxiv

Modern video generators still struggle with complex physical dynamics, often falling short of physical realism. Existing approaches address this using external verifiers or additional training on augmented data, which is computationally expensive and still limited in capturing fine-grained motion. In this work, we present self-refining video sampling, a simple method that uses a pre-trained video generator trained on large-scale datasets as its own self-refiner. By interpreting the generator as a denoising autoencoder, we enable iterative inner-loop refinement at inference time without any external verifier or additional training. We further introduce an uncertainty-aware refinement strategy that selectively refines regions based on self-consistency, which prevents artifacts caused by over-refinement. Experiments on state-of-the-art video generators demonstrate significant improvements in motion coherence and physics alignment, achieving over 70% human preference compared to the default sampler and guidance-based sampler.

📄 PDF Abstract BibTeX arXiv:2601.18577

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hierarchical Frequency-based Upsampling and Refining for Compressed Video Quality Enhancement

2024-03-18 · Qianyu Zhang, Bolun Zheng, Xinying Chen, Quan Chen 외

Video compression artifacts arise due to the quantization operation in the frequency domain. The goal of video quality enhancement is to reduce compression artifacts and reconstruct a visually-pleasant result. In this wo…

QuantizationVideo Compression

VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

2024-03-08 · Yabo Zhang, Yuxiang Wei, Xianhui Lin, Zheng Hui 외

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality an…

Video Generation

TimeRefine: Temporal Grounding with Time Refining Video LLM

2024-12-12 · Xizi Wang, Feng Cheng, Ziyang Wang, Huiyu Wang 외

Video temporal grounding aims to localize relevant temporal boundaries in a video given a textual prompt. Recent work has focused on enabling Video LLMs to perform video temporal grounding via next-token prediction of te…

Temporal Localization

ReMOTS: Self-Supervised Refining Multi-Object Tracking and Segmentation

2020-07-07 · Fan Yang, Xin Chang, Chenyu Dang, Ziqiang Zheng 외

We aim to improve the performance of Multiple Object Tracking and Segmentation (MOTS) by refinement. However, it remains challenging for refining MOTS results, which could be attributed to that appearance features are no…

Multi-Object TrackingMulti-Object Tracking and SegmentationMultiple Object TrackingObject+1

CRVOS: Clue Refining Network for Video Object Segmentation

2020-02-10 · Suhwan Cho, MyeongAh Cho, Tae-young Chung, Heansung Lee 외

The encoder-decoder based methods for semi-supervised video object segmentation (Semi-VOS) have received extensive attention due to their superior performances. However, most of them have complex intermediate networks wh…

DecoderObjectSegmentationSemantic Segmentation+4