DTrOCR: Decoder-only Transformer for Optical Character Recognition
Typical text recognition methods rely on an encoder-decoder structure, in which the encoder extracts features from an image, and the decoder produces recognized text from these features. In this study, we propose a simpler and more effective method for text recognition, known as the Decoder-only Transformer for Optical Character Recognition (DTrOCR). This method uses a decoder-only Transformer to take advantage of a generative language model that is pre-trained on a large corpus. We examined whether a generative language model that has been successful in natural language processing can also be effective for text recognition in computer vision. Our experiments demonstrated that DTrOCR outperforms current state-of-the-art methods by a large margin in the recognition of printed, handwritten, and scene text in both English and Chinese.
Code (1)
Tasks
DecoderHandwritten Text RecognitionLanguage ModelingLanguage ModellingOptical Character RecognitionOptical Character Recognition (OCR)Scene Text RecognitionTask 2Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Transformer-Based Approach for Diagnosing Fault Cases in Optical Fiber Amplifiers
A transformer-based deep learning approach is presented that enables the diagnosis of fault cases in optical fiber amplifiers using condition-based monitoring time series data. The model, Inverse Triple-Aspect Self-Atten…
DecoderTime SeriesAmodal Optical Flow
Optical flow estimation is very challenging in situations with transparent or occluded objects. In this work, we address these challenges at the task level by introducing Amodal Optical Flow, which integrates optical flo…
DecoderOptical Flow EstimationPanoptic TrackingOptoGPT: A Foundation Model for Inverse Design in Optical Multilayer Thin Film Structures
Optical multilayer thin film structures have been widely used in numerous photonic applications. However, existing inverse design methods have many drawbacks because they either fail to quickly adapt to different design …
Computational EfficiencyDecoderMED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
In this paper, we present an end-to-end trainable unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in video. The presented Multiscale Encoder-Decoder Video Transformer (MED-VT) use…
Action SegmentationDecoderOptical Flow EstimationSegmentation+5Handwritten Text Recognition for Low Resource Languages
Despite considerable progress in handwritten text recognition, paragraph-level handwritten text recognition, especially in low-resource languages, such as Hindi, Urdu and similar scripts, remains a challenging problem. T…
Handwritten Text Recognition