paper-with-me

Papers

What Makes AI Applications Acceptable or Unacceptable? A Predictive Moral Framework

2025-08-26 · Kimmo Eriksson, Simon Karlsson, Irina Vartanova, Pontus Strimling arxiv

As artificial intelligence rapidly transforms society, developers and policymakers struggle to anticipate which applications will face public moral resistance. We propose that these judgments are not idiosyncratic but systematic and predictable. In a large, preregistered study (N = 587, U.S. representative sample), we used a comprehensive taxonomy of 100 AI applications spanning personal and organizational contexts-including both functional uses and the moral treatment of AI itself. In participants' collective judgment, applications ranged from highly unacceptable to fully acceptable. We found this variation was strongly predictable: five core moral qualities-perceived risk, benefit, dishonesty, unnaturalness, and reduced accountability-collectively explained over 90% of the variance in acceptability ratings. The framework demonstrated strong predictive power across all domains and successfully predicted individual-level judgments for held-out applications. These findings reveal that a structured moral psychology underlies public evaluation of new technologies, offering a powerful tool for anticipating public resistance and guiding responsible innovation in AI.

📄 PDF Abstract BibTeX arXiv:2508.19317

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Espresso: Robust Concept Filtering in Text-to-Image Models

2024-04-30 · Anudeep Das, Vasisht Duddu, Rui Zhang, N. Asokan

Diffusion based text-to-image models are trained on large datasets scraped from the Internet, potentially containing unacceptable concepts (e.g., copyright-infringing or unsafe). We need concept removal techniques (CRTs)…

Retweet communities reveal the main sources of hate speech

2021-05-31 · Bojan Evkoski, Andraz Pelicon, Igor Mozetic, Nikola Ljubesic 외

We address a challenging problem of identifying main sources of hate speech on Twitter. On one hand, we carefully annotate a large set of tweets for hate speech, and deploy advanced deep learning to produce high quality …

Do Concept Replacement Techniques Really Erase Unacceptable Concepts?

2025-06-10 · Anudeep Das, Gurjot Singh, Prach Chantasantitam, N. Asokan

Generative models, particularly diffusion-based text-to-image (T2I) models, have demonstrated astounding success. However, aligning them to avoid generating content with unacceptable concepts (e.g., offensive or copyrigh…

Legal Framework, Dataset and Annotation Schema for Socially Unacceptable Online Discourse Practices in Slovene

2017-08-01 · WS 2017 8 · Darja Fi{\v{s}}er, Toma{\v{z}} Erjavec, Nikola Ljube{\v{s}}i{\'c}

In this paper we present the legal framework, dataset and annotation schema of socially unacceptable discourse practices on social networking platforms in Slovenia. On this basis we aim to train an automatic identificati…

General Classification

Detecting Community Sensitive Norm Violations in Online Conversations

2021-10-09 · Findings (EMNLP) 2021 11 · Chan Young Park, Julia Mendelsohn, Karthik Radhakrishnan, Kinjal Jain 외

Online platforms and communities establish their own norms that govern what behavior is acceptable within the community. Substantial effort in NLP has focused on identifying unacceptable behaviors and, recently, on forec…