paper-with-me

홈 › Papers

Joint Energy-based Detection and Classificationon of Multilingual Text Lines

2014-07-23 · Igor Milevskiy, Yuri Boykov

This paper proposes a new hierarchical MDL-based model for a joint detection and classification of multilingual text lines in im- ages taken by hand-held cameras. The majority of related text detec- tion methods assume alphabet-based writing in a single language, e.g. in Latin. They use simple clustering heuristics specific to such texts: prox- imity between letters within one line, larger distance between separate lines, etc. We are interested in a significantly more ambiguous problem where images combine alphabet and logographic characters from multiple languages and typographic rules vary a lot (e.g. English, Korean, and Chinese). Complexity of detecting and classifying text lines in multiple languages calls for a more principled approach based on information- theoretic principles. Our new MDL model includes data costs combining geometric errors with classification likelihoods and a hierarchical sparsity term based on label costs. This energy model can be efficiently minimized by fusion moves. We demonstrate robustness of the proposed algorithm on a large new database of multilingual text images collected in the pub- lic transit system of Seoul.

📄 PDF Abstract BibTeX arXiv:1407.6082

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringGeneral Classification

Methods 이 논문이 사용한 방법론

MDL Minimum Description Length provides a criterion for the selection of models, regardless of their complexity, without the restrictive assumption that the data form a sample…

Similar Papers 제목 키워드 기반

UPV at CheckThat! 2021: Mitigating Cultural Differences for Identifying Multilingual Check-worthy Claims

2021-09-19 · Ipek Baris Schlicht, Angel Felipe Magnossão de Paula, Paolo Rosso

Identifying check-worthy claims is often the first step of automated fact-checking systems. Tackling this task in a multilingual setting has been understudied. Encoding inputs with multilingual text representations could…

Fact CheckingLanguage Identification

MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

2026-05-07 · Sofia Callejas, Nahuel Gomez, Catherine Pelachaud, Brian Ravenet 외 arxiv

Laughter is a social non-vocalization that is universal across cultures and languages, and is crucial for human communication, including social bonding and communication signaling. However, detecting laughter in audio is…

Anomaly Detection

ANDES at SemEval-2020 Task 12: A jointly-trained BERT multilingual model for offensive language detection

2020-08-13 · SEMEVAL 2020 · Juan Manuel Pérez, Aymé Arango, Franco Luque

This paper describes our participation in SemEval-2020 Task 12: Multilingual Offensive Language Detection. We jointly-trained a single model by fine-tuning Multilingual BERT to tackle the task across all the proposed lan…

UTCNN: a Deep Learning Model of Stance Classificationon on Social Media Text

2016-11-11 · Wei-Fan Chen, Lun-Wei Ku

Most neural network models for document classification on social media focus on text infor-mation to the neglect of other information on these platforms. In this paper, we classify post stance on social media channels an…

Document Classification

Multi-Label Out-of-Distribution Detection with Spectral Normalized Joint Energy

2024-05-08 · Yihan Mei, Xinyu Wang, Dell Zhang, Xiaoling Wang

In today's interconnected world, achieving reliable out-of-distribution (OOD) detection poses a significant challenge for machine learning models. While numerous studies have introduced improved approaches for multi-clas…

Out-of-Distribution DetectionOut of Distribution (OOD) Detection