paper-with-me

홈 › Papers

Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts

2023-09-25 · Thibault Clérice

In this study, we propose to evaluate the use of deep learning methods for semantic classification at the sentence level to accelerate the process of corpus building in the field of humanities and linguistics, a traditional and time-consuming task. We introduce a novel corpus comprising around 2500 sentences spanning from 300 BCE to 900 CE including sexual semantics (medical, erotica, etc.). We evaluate various sentence classification approaches and different input embedding layers, and show that all consistently outperform simple token-based searches. We explore the integration of idiolectal and sociolectal metadata embeddings (centuries, author, type of writing), but find that it leads to overfitting. Our results demonstrate the effectiveness of this approach, achieving high precision and true positive rates (TPR) of respectively 70.60% and 86.33% using HAN. We evaluate the impact of the dataset size on the model performances (420 instead of 2013), and show that, while our models perform worse, they still offer a high enough precision and TPR, even without MLM, respectively 69% and 51%. Given the result, we provide an analysis of the attention mechanism as a supporting added value for humanists in order to produce more data.

📄 PDF Abstract BibTeX arXiv:2309.14974

Code (1)

lascivaroma/seligator 공식 구현 pytorch

Tasks

SentenceSentence Classification

Similar Papers 제목 키워드 기반

Detecting sexually explicit content in the context of the child sexual abuse materials (CSAM): end-to-end classifiers and region-based networks

2024-06-20 · Weronika Gutfeter, Joanna Gajewska, Andrzej Pacut

Child sexual abuse materials (CSAM) pose a significant threat to the safety and well-being of children worldwide. Detecting and preventing the distribution of such materials is a critical task for law enforcement agencie…

Human Detection

Queer People are People First: Deconstructing Sexual Identity Stereotypes in Large Language Models

2023-06-30 · Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, Emma Strubell

Large Language Models (LLMs) are trained primarily on minimally processed web text, which exhibits the same wide range of social biases held by the humans who created that content. Consequently, text generated by LLMs ca…

Sentence

Detecting Harmful Content On Online Platforms: What Platforms Need Vs. Where Research Efforts Go

2021-02-27 · Arnav Arora, Preslav Nakov, Momchil Hardalov, Sheikh Muhammad Sarwar 외

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence…

Abusive LanguageMisinformation

NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

2025-10-17 · Yitong Sun, Yao Huang, Ruochen Zhang, Huanran Chen 외 arxiv

Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with implicit sexual prompts. Unlike explicit…

Text-to-Image Generation

Detecting Offensive Content in Open-domain Conversations using Two Stage Semi-supervision

2018-11-30 · Chandra Khatri, Behnam Hedayatnia, Rahul Goel, Anushree Venkatesh 외

As open-ended human-chatbot interaction becomes commonplace, sensitive content detection gains importance. In this work, we propose a two stage semi-supervised approach to bootstrap large-scale data for automatic sensiti…

Chatbot