paper-with-me

홈 › Papers

Novel Keyword Extraction and Language Detection Approaches

2020-09-24 · Malgorzata Pikies, Andronicus Riyono, Junade Ali

Fuzzy string matching and language classification are important tools in Natural Language Processing pipelines, this paper provides advances in both areas. We propose a fast novel approach to string tokenisation for fuzzy language matching and experimentally demonstrate an 83.6% decrease in processing time with an estimated improvement in recall of 3.1% at the cost of a 2.6% decrease in precision. This approach is able to work even where keywords are subdivided into multiple words, without needing to scan character-to-character. So far there has been little work considering using metadata to enhance language classification algorithms. We provide observational data and find the Accept-Language header is 14% more likely to match the classification than the IP Address.

📄 PDF Abstract BibTeX arXiv:2009.11832

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationKeyword Extraction

Similar Papers 제목 키워드 기반

Out of Thin Air: Is Zero-Shot Cross-Lingual Keyword Detection Better Than Unsupervised?

2022-02-14 · LREC 2022 6 · Boshko Koloski, Senja Pollak, Blaž Škrlj, Matej Martinc

Keyword extraction is the task of retrieving words that are essential to the content of a given document. Researchers proposed various approaches to tackle this problem. At the top-most level, approaches are divided into…

Keyword ExtractionPretrained Multilingual Language Models

PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection

2025-10-10 · Rakib Hossan, Shubhashis Roy Dipta arxiv

The BLP-2025 Task 1A requires Bengali hate speech classification into six categories. Traditional supervised approaches need extensive labeled datasets that are expensive for low-resource languages. We developed PromptGu…

Hate Speech DetectionKeyword Extraction

Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI

2026-07-10 · Miguel Arana-Catania, Catherine Conisbee, Matthew Kidd arxiv

Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" proj…

Keyword Extraction

Exploring Linguistically-Lightweight Keyword Extraction Techniques for Indexing News Articles in a Multilingual Set-up

2021-04-01 · EACL (Hackashop) 2021 4 · Jakub Piskorski, Nicolas Stefanovitch, Guillaume Jacquet, Aldo Podavini

This paper presents a study of state-of-the-art unsupervised and linguistically unsophisticated keyword extraction algorithms, based on statistic-, graph-, and embedding-based approaches, including, i.a., Total Keyword F…

ArticlesKeyword Extraction

Thematic context vector association based on event uncertainty for Twitter

2023-04-04 · Vaibhav Khatavkar, Swapnil Mane, Parag Kulkarni

Keyword extraction is a crucial process in text mining. The extraction of keywords with respective contextual events in Twitter data is a big challenge. The challenging issues are mainly because of the informality in the…

Keyword ExtractionSarcasm Detection