paper-with-me

Papers

TransLPRNet: Lite Vision-Language Network for Single/Dual-line Chinese License Plate Recognition

2025-07-23 · Guangzhu Xu, Zhi Ke, Pengcheng Zuo, Bangjun Lei arxiv

License plate recognition in open environments is widely applicable across various domains; however, the diversity of license plate types and imaging conditions presents significant challenges. To address the limitations encountered by CNN and CRNN-based approaches in license plate recognition, this paper proposes a unified solution that integrates a lightweight visual encoder with a text decoder, within a pre-training framework tailored for single and double-line Chinese license plates. To mitigate the scarcity of double-line license plate datasets, we constructed a single/double-line license plate dataset by synthesizing images, applying texture mapping onto real scenes, and blending them with authentic license plate images. Furthermore, to enhance the system's recognition accuracy, we introduce a perspective correction network (PTN) that employs license plate corner coordinate regression as an implicit variable, supervised by license plate view classification information. This network offers improved stability, interpretability, and low annotation costs. The proposed algorithm achieves an average recognition accuracy of 99.34% on the corrected CCPD test set under coarse localization disturbance. When evaluated under fine localization disturbance, the accuracy further improves to 99.58%. On the double-line license plate test set, it achieves an average recognition accuracy of 98.70%, with processing speeds reaching up to 167 frames per second, indicating strong practical applicability.

📄 PDF Abstract BibTeX arXiv:2507.17335

Code (0)

등록된 구현이 없습니다.

Tasks

License Plate Recognition

Similar Papers 제목 키워드 기반

Modulating early visual processing by language

2017-07-02 · NeurIPS 2017 12 · Harm de Vries, Florian Strub, Jérémie Mary, Hugo Larochelle 외

It is commonly assumed that language refers to high-level visual concepts while leaving low-level visual processing unaffected. This view dominates the current literature in computational models for language-vision tasks…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

MatMMExtract: An Open-Source Pipeline for Panel-Level Extraction of Grounded Image-Text Pairs from Materials Science Literature

2026-06-29 · Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari hf

The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale. The core difficulty is structural: most scientific figure…

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

2026-06-02 · Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Luyao Ye 외 arxiv

When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author is trying to communicate. Standard instruction tuning entangles a pos…

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

2025-12-02 · Kun Yuan, Min Woo Sun, Zhen Chen, Alejandro Lozano 외 arxiv

There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. However, current biomedical vision-language pr…

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

2024-06-16 · Wenyan Li, Xinyu Zhang, Jiaang Li, Qiwei Peng 외

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQ…

DiversityMultiple-choiceQuestion AnsweringVisual Question Answering (VQA)