paper-with-me

Papers

Multimodal Semi-Supervised Learning for Text Recognition

2022-05-08 · Aviad Aberdam, Roy Ganz, Shai Mazor, Ron Litman

Until recently, the number of public real-world text images was insufficient for training scene text recognizers. Therefore, most modern training methods rely on synthetic data and operate in a fully supervised manner. Nevertheless, the amount of public real-world text images has increased significantly lately, including a great deal of unlabeled data. Leveraging these resources requires semi-supervised approaches; however, the few existing methods do not account for vision-language multimodality structure and therefore suboptimal for state-of-the-art multimodal architectures. To bridge this gap, we present semi-supervised learning for multimodal text recognizers (SemiMTR) that leverages unlabeled data at each modality training phase. Notably, our method refrains from extra training stages and maintains the current three-stage multimodal training procedure. Our algorithm starts by pretraining the vision model through a single-stage training that unifies self-supervised learning with supervised training. More specifically, we extend an existing visual representation learning algorithm and propose the first contrastive-based method for scene text recognition. After pretraining the language model on a text corpus, we fine-tune the entire network via a sequential, character-level, consistency regularization between weakly and strongly augmented views of text images. In a novel setup, consistency is enforced on each modality separately. Extensive experiments validate that our method outperforms the current training schemes and achieves state-of-the-art results on multiple scene text recognition benchmarks.

📄 PDF Abstract BibTeX arXiv:2205.03873

Code (2)

amazon-research/semimtr-text-recognition 공식 구현 pytorch
amazon-science/semimtr-text-recognition pytorch

Tasks

Language ModellingRepresentation LearningScene Text RecognitionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment Recognition

2025-01-01 · CVPR 2025 1 · Wuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng Yang

Multimodal sentiment analysis has attracted extensive research attention as increasing users share images and texts to express their emotions and opinions on social media. Collecting large amounts of labeled sentimen…

Multimodal Sentiment AnalysisPseudo LabelPseudo Label FilteringSentiment Analysis

Learning from Temporal Gradient for Semi-supervised Action Recognition

2021-11-25 · CVPR 2022 1 · Junfei Xiao, Longlong Jing, Lin Zhang, Ju He 외

Semi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-bas…

Action RecognitionTemporal Action Localization

MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning

2023-04-18 · Zheng Lian, Haiyang Sun, Licai Sun, Kang Chen 외

The first Multimodal Emotion Recognition Challenge (MER 2023) was successfully held at ACM Multimedia. The challenge focuses on system robustness and consists of three distinct tracks: (1) MER-MULTI, where participants a…

Emotion RecognitionMulti-Label LearningMultimodal Emotion Recognition

SemiMemes: A Semi-supervised Learning Approach for Multimodal Memes Analysis

2023-03-31 · Pham Thai Hoang Tung, Nguyen Tan Viet, Ngo Tien Anh, Phan Duy Hung

The prevalence of memes on social media has created the need to sentiment analyze their underlying meanings for censoring harmful content. Meme censoring systems by machine learning raise the need for a semi-supervised l…

Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples

2024-08-23 · Qi Fan, Yutong Li, Yi Xin, Xinyu Cheng 외

The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions for the Semi-Supervised Learning Sub-Chal…

Contrastive LearningEmotion RecognitionMultimodal Emotion Recognition