paper-with-me

홈 › Papers

MENTOR: Multilingual tExt detectioN TOward leaRning by analogy

2024-03-12 · Hsin-Ju Lin, Tsu-Chun Chung, Ching-Chun Hsiao, Pin-Yu Chen, Wei-Chen Chiu, Ching-Chun Huang

Text detection is frequently used in vision-based mobile robots when they need to interpret texts in their surroundings to perform a given task. For instance, delivery robots in multilingual cities need to be capable of doing multilingual text detection so that the robots can read traffic signs and road markings. Moreover, the target languages change from region to region, implying the need of efficiently re-training the models to recognize the novel/new languages. However, collecting and labeling training data for novel languages are cumbersome, and the efforts to re-train an existing/trained text detector are considerable. Even worse, such a routine would repeat whenever a novel language appears. This motivates us to propose a new problem setting for tackling the aforementioned challenges in a more efficient way: "We ask for a generalizable multilingual text detection framework to detect and identify both seen and unseen language regions inside scene images without the requirement of collecting supervised training data for unseen languages as well as model re-training". To this end, we propose "MENTOR", the first work to realize a learning strategy between zero-shot learning and few-shot learning for multilingual scene text detection.

📄 PDF Abstract BibTeX arXiv:2403.07286

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningScene Text DetectionText DetectionZero-Shot Learning

Similar Papers 제목 키워드 기반

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

2026-01-23 · Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat arxiv

Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that provide reflection and guidance. Existi…

Question Answering

Multilingual Culture-Independent Word Analogy Datasets

2019-11-22 · LREC 2020 5 · Matej Ulčar, Kristiina Vaik, Jessica Lindström, Milda Dailidėnaitė 외

In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-dimensional numeric space. To compare the …

Cultural Vocal Bursts Intensity PredictionWord Embeddings

YT-30M: A multi-lingual multi-category dataset of YouTube comments

2024-12-04 · Hridoy Sankar Dutta

This paper introduces two large-scale multilingual comment datasets, YT-30M (and YT-100K) from YouTube. The analysis in this paper is performed on a smaller sample (YT-100K) of YT-30M. Both the datasets: YT-30M (full) an…

Mentor3AD: Feature Reconstruction-based 3D Anomaly Detection via Multi-modality Mentor Learning

2025-05-27 · Jinbao Wang, Hanzhe Liang, Can Gao, Chenxi Hu 외

Multimodal feature reconstruction is a promising approach for 3D anomaly detection, leveraging the complementary information from dual modalities. We further advance this paradigm by utilizing multi-modal mentor learning…

3D Anomaly DetectionAnomaly Detection

AOMD: An Analogy-aware Approach to Offensive Meme Detection on Social Media

2021-06-21 · Lanyu Shang, Yang Zhang, Yuheng Zha, Yingxi Chen 외

This paper focuses on an important problem of detecting offensive analogy meme on online social media where the visual content and the texts/captions of the meme together make an analogy to convey the offensive informati…