paper-with-me

Papers

Text-DIAE: A Self-Supervised Degradation Invariant Autoencoders for Text Recognition and Document Enhancement

2022-03-09 · Mohamed Ali Souibgui, Sanket Biswas, Andres Mafla, Ali Furkan Biten, Alicia Fornés, Yousri Kessentini, Josep Lladós, Lluis Gomez, Dimosthenis Karatzas

In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext tasks as learning objectives to be optimized during pre-training without the usage of labeled data. Each of the pretext objectives is specifically tailored for the final downstream tasks. We conduct several ablation experiments that confirm the design choice of the selected pretext tasks. Importantly, the proposed model does not exhibit limitations of previous state-of-the-art methods based on contrastive losses, while at the same time requiring substantially fewer data samples to converge. Finally, we demonstrate that our method surpasses the state-of-the-art in existing supervised and self-supervised settings in handwritten and scene text recognition and document image enhancement. Our code and trained models will be made publicly available at~\url{ http://Upon_Acceptance}.

📄 PDF Abstract BibTeX arXiv:2203.04814

Code (1)

dali92002/SSL-OCR 공식 구현 pytorch

Tasks

Document EnhancementImage EnhancementScene Text Recognition

Similar Papers 제목 키워드 기반

Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation

2025-03-20 · Jiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie 외

In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to enhance the sharpness and generalizatio…

Depth EstimationImage ReconstructionMonocular Depth EstimationZero-shot Generalization

DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models

2024-10-31 · Heng-Jui Chang, Hongyu Gong, Changhan Wang, James Glass 외

Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This …

DecoderResynthesisSpeech Tokenization

Self-Supervised Learning of Pretext-Invariant Representations

2019-12-04 · CVPR 2020 6 · Ishan Misra, Laurens van der Maaten

The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations for a large training set of images. Many …

Contrastive Learningobject-detectionObject DetectionRepresentation Learning+3

DRIFT: Drift-Resilient Invariant-Feature Transformer for DGA Detection

2026-05-11 · Chaeyoung Lee, Chaeri Jung, Seonghoon Jeong arxiv

Domain Generation Algorithms (DGAs) evolve continuously to evade botnet detection, posing a persistent challenge for dependable network defense. While deep learning-based detectors achieve strong performance under static…

TUKE at MediaEval 2015 QUESST

2015-09-14 · MediaEval 2015 Workshop 2015 9 · Jozef Vavrek, Peter Viszlay, Martin Lojka, Matúš Pleva 외

In this paper, we present our retrieving system for QUery by Example Search on Speech Task (QUESST), comprising the posteriorgram-based modeling approach along with the weighted fast sequential dynamic time warping algor…

Dynamic Time WarpingKeyword Spotting