paper-with-me

홈 › Papers

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal

2026-03-23 · Qingdong He, Chaoyi Wang, Peng Tang, Yifan Yang, Xiaobin Hu arxiv

Video subtitle removal aims to distinguish text overlays from background content while preserving temporal coherence. Existing diffusion-based methods necessitate explicit mask sequences during both training and inference phases, which restricts their practical deployment. In this paper, we present CLEAR (Context-aware Learning for End-to-end Adaptive Video Subtitle Removal), a mask-free framework that achieves truly end-to-end inference through context-aware adaptive learning. Our two-stage design decouples prior extraction from generative refinement: Stage I learns disentangled subtitle representations via self-supervised orthogonality constraints on dual encoders, while Stage II employs LoRA-based adaptation with generation feedback for dynamic context adjustment. Notably, our method only requires 0.77% of the parameters of the base diffusion model for training. On Chinese subtitle benchmarks, CLEAR outperforms mask-dependent baselines by + 6.77dB PSNR and -74.7% VFID, while demonstrating superior zero-shot generalization across six languages (English, Korean, French, Japanese, Russian, German), a performance enabled by our generation-driven feedback mechanism that ensures robust subtitle removal without ground-truth masks during inference.

📄 PDF Abstract BibTeX arXiv:2603.21901

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

SAFE-DiT: Semantics-Aware Fast-path Execution for High-Resolution Diffusion Transformers

2026-06-28 · Xuanhua Yin, Yuxuan Jia, Chuanzhi Xu, Weidong Cai arxiv

High-resolution Diffusion Transformer (DiT) inference contains substantial spatial redundancy, but many spatially adaptive implementations encode regional computation as attention masks, which can inadvertently move scal…

UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training

2026-05-26 · Keqi Deng, Shaoshi Ling, Ruchao Fan, Jinyu Li arxiv

Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse attention alleviates this by loading only a small fraction of the KV ca…

Speech Recognition

Nuclear Instance Segmentation using a Proposal-Free Spatially Aware Deep Learning Framework

2019-08-27 · Navid Alemi Koohbanani, Mostafa Jahanifar, Ali Gooya, Nasir Rajpoot

Nuclear segmentation in histology images is a challenging task due to significant variations in the shape and appearance of nuclei. One of the main hurdles in nuclear instance segmentation is overlapping nuclei where a s…

ClusteringInstance SegmentationNuclear SegmentationSegmentation+1

Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models

2025-12-22 · Tongyuan Miao, Gary Huang, Kai Jun Han, Annie Jiang arxiv

Diffusion Large Language Models (DLLMs) enable fully parallel token decoding but often remain impractical at inference time due to the many denoising iterations required to refine an information-free, fully masked initia…

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

2026-08-01 · Marcel Plocher, Bernhard Schölkopf, Andreas Geiger, Gege Gao arxiv

The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumption is particularly questionable for continuous masked generators, wh…