paper-with-me

Papers

Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks

2025-12-02 · Matthew Dutson, Nathan Labiosa, Yin Li, Mohit Gupta arxiv

When applied sequentially to video, frame-based networks often exhibit temporal inconsistency - for example, outputs that flicker between frames. This problem is amplified when the network inputs contain time-varying corruptions. In this work, we introduce a general approach for adapting frame-based models for stable and robust inference on video. We describe a class of stability adapters that can be inserted into virtually any architecture and a resource-efficient training process that can be performed with a frozen base network. We introduce a unified conceptual framework for describing temporal stability and corruption robustness, centered on a proposed accuracy-stability-robustness loss. By analyzing the theoretical properties of this loss, we identify the conditions where it produces well-behaved stabilizer training. Our experiments validate our approach on several vision tasks including denoising (NAFNet), image enhancement (HDRNet), monocular depth (Depth Anything v2), and semantic segmentation (DeepLabv3+). Our method improves temporal stability and robustness against a range of image corruptions (including compression artifacts, noise, and adverse weather), while preserving or improving the quality of predictions.

📄 PDF Abstract BibTeX arXiv:2512.03014

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationImage Enhancement

Similar Papers 제목 키워드 기반

MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model

2024-05-30 · Muyao Niu, Xiaodong Cun, Xintao Wang, Yong Zhang 외

We present MOFA-Video, an advanced controllable image animation method that generates video from the given image using various additional controllable signals (such as human landmarks reference, manual trajectories, and …

Image AnimationVideo Generation

Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals

2026-05-30 · Shenghao Ding arxiv

This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around …

Language-Universal Adapter Learning with Knowledge Distillation for End-to-End Multilingual Speech Recognition

2023-02-28 · Zhijie Shen, Wu Guo, Bin Gu

In this paper, we propose a language-universal adapter learning framework based on a pre-trained model for end-to-end multilingual automatic speech recognition (ASR). For acoustic modeling, the wav2vec 2.0 pre-trained mo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillationspeech-recognition+1

Universal Few-Shot Spatial Control for Diffusion Models

2025-09-09 · Kiet T. Nguyen, Chanhyuk Lee, Donggyun Kim, Dong Hoon Lee 외 arxiv

Spatial conditioning in pretrained text-to-image diffusion models has significantly improved fine-grained control over the structure of generated images. However, existing control adapters exhibit limited adaptability an…

One Model to Translate Them All: Universal Any-to-Any Translation for Heterogeneous Collaborative Perception

2026-05-18 · Yang Li, Weize Li, Quan Yuan, Congzhang Shao 외 arxiv

By sharing intermediate features, collaborative perception extends each agent's sensing beyond standalone limits, but real-world feature modality heterogeneity remains a key barrier to effective fusion. Most existing met…