paper-with-me

Papers

Detection of adversarial intent in Human-AI teams using LLMs

2026-03-21 · Abed K. Musaffar, Ambuj Singh, Francesco Bullo arxiv

Large language models (LLMs) are increasingly deployed in human-AI teams as support agents for complex tasks such as information retrieval, programming, and decision-making assistance. While these agents' autonomy and contextual knowledge enables them to be useful, it also exposes them to a broad range of attacks, including data poisoning, prompt injection, and even prompt engineering. Through these attack vectors, malicious actors can manipulate an LLM agent to provide harmful information, potentially manipulating human agents to make harmful decisions. While prior work has focused on LLMs as attack targets or adversarial actors, this paper studies their potential role as defensive supervisors within mixed human-AI teams. Using a dataset consisting of multi-party conversations and decisions for a real human-AI team over a 25 round horizon, we formulate the problem of malicious behavior detection from interaction traces. We find that LLMs are capable of identifying malicious behavior in real-time, and without task-specific information, indicating the potential for task-agnostic defense. Moreover, we find that the malicious behavior of interest is not easily identified using simple heuristics, further suggesting the introduction of LLM defenders could render human teams more robust to certain classes of attack.

📄 PDF Abstract BibTeX arXiv:2603.20976

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalPrompt Engineering

Similar Papers 제목 키워드 기반

ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations

2025-07-17 · Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira 외 arxiv

The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user inten…

M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text

2025-11-14 · Salima Lamsiyah, Saad Ezzini, Abdelkader El Mahdaouy, Hamza Alami 외 arxiv

The generation of highly fluent text by Large Language Models (LLMs) poses a significant challenge to information integrity and academic research. In this paper, we introduce the Multi-Domain Detection of AI-Generated Te…

Binary Classification

MALicious INTent Dataset and Inoculating LLMs for Enhanced Disinformation Detection

2026-03-15 · Arkadiusz Modzelewski, Witold Sosnowski, Eleni Papadopulos, Elisa Sartori 외 arxiv

The intentional creation and spread of disinformation poses a significant threat to public discourse. However, existing English datasets and research rarely address the intentionality behind the disinformation. This work…

Intent Classification

GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge

2024-12-24 · Shammur Absar Chowdhury, Hind Almerekhi, Mucahid Kutlu, Kaan Efe Keles 외

This paper presents a comprehensive overview of the first edition of the Academic Essay Authenticity Challenge, organized as part of the GenAI Content Detection shared tasks collocated with COLING 2025. This challenge fo…

Task 2

Controllable Conversational Theme Detection Track at DSTC 12

2025-08-26 · Igor Shalyminov, Hang Su, Jake Vincent, Siffi Singh 외 arxiv

Conversational analytics has been on the forefront of transformation driven by the advances in Speech and Natural Language Processing techniques. Rapid adoption of Large Language Models (LLMs) in the analytics field has …

Intent Detection