paper-with-me

홈 › Papers

Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse

2024-12-23 · Anna Kołos, Katarzyna Lorenc, Emilia Wiśnios, Agnieszka Karlińska

The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. We present forePLay, a novel Polish language dataset for erotic content detection, featuring over 24k annotated sentences with a multidimensional taxonomy encompassing ambiguity, violence, and social unacceptability dimensions. Our comprehensive evaluation demonstrates that specialized Polish language models achieve superior performance compared to multilingual alternatives, with transformer-based architectures showing particular strength in handling imbalanced categories. The dataset and accompanying analysis establish essential frameworks for developing linguistically-aware content moderation systems, while highlighting critical considerations for extending such capabilities to morphologically complex languages.

📄 PDF Abstract BibTeX arXiv:2412.17533

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accuracy of the Uzbek stop words detection: a case study on "School corpus"

2022-09-15 · Khabibulla Madatov, Shukurla Bekchanov, Jernej Vičič

Stop words are very important for information retrieval and text analysis investigation tasks of natural language processing. Current work presents a method to evaluate the quality of a list of stop words aimed at automa…

Information RetrievalRetrievalSentence

Building a Public Domain Voice Database for Odia

2022-08-16 · WWW '22: Companion Proceedings of the Web Conference 2022 8 · Subhashish Panigrahi

Projects like Mozilla Common Voice were born to address the challenges of unavailability of voice data or the high cost of available data for use in speech technology such as Automatic Speech Recognition (ASR) research a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Samāsa-Kartā: An Online Tool for Producing Compound Words using IndoWordNet

2016-01-01 · GWC 2016 1 · Hanumant Redkar, Nilesh Joshi, Sandhya Singh, Irawati Kulkarni 외

Samāsa or compounds are a regular feature of Indian Languages. They are also found in other languages like German, Italian, French, Russian, Spanish, etc. Compound word is constructed from two or more words to form a sin…

Morphological Analysis

Investigating the structure of emotions by analyzing similarity and association of emotion words

2026-02-06 · Fumitaka Iwaki, Tatsuji Takahashi arxiv

In the field of natural language processing, some studies have attempted sentiment analysis on text by handling emotions as explanatory or response variables. One of the most popular emotion models used in this context i…

Community DetectionSentiment Analysis

Tuiteamos o pongamos un tuit? Investigating the Social Constraints of Loanword Integration in Spanish Social Media

2021-01-16 · SCiL 2021 2 · Ian Stewart, Diyi Yang, Jacob Eisenstein

Speakers of non-English languages often adopt loanwords from English to express new or unusual concepts. While these loanwords may be borrowed unchanged, speakers may also integrate the words to fit the constraints of th…