paper-with-me

홈 › Papers

Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design

2025-07-15 · Richard M. Charles, James H. Curry, Richard B. Charles arxiv

The integration of Large Language Models (LLMs) in K--12 education offers both transformative opportunities and emerging risks. This study explores how students may Trojanize prompts to elicit unsafe or unintended outputs from LLMs, bypassing established content moderation systems with safety guardrils. Through a systematic experiment involving simulated K--12 queries and multi-turn dialogues, we expose key vulnerabilities in GPT-3.5 and GPT-4. This paper presents our experimental design, detailed findings, and a prototype tool, TrojanPromptGuard (TPG), to automatically detect and mitigate Trojanized educational prompts. These insights aim to inform both AI safety researchers and educational technologists on the safe deployment of LLMs for educators.

📄 PDF Abstract BibTeX arXiv:2507.14207

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI4OPT: AI Institute for Advances in Optimization

2023-07-05 · Pascal Van Hentenryck, Kevin Dalmeijer

This article is a short introduction to AI4OPT, the NSF AI Institute for Advances in Optimization. AI4OPT fuses AI and Optimization, inspired by end-use cases in supply chains, energy systems, chip design and manufacturi…

Philosophy

LLM Prompt Evaluation for Educational Applications

2026-01-22 · Langdon Holmes, Adam Coscia, Scott Crossley, Joon Suh Choi 외 arxiv

As large language models (LLMs) become increasingly common in educational applications, there is a growing need for evidence-based methods to design and evaluate LLM prompts that produce personalized and pedagogically al…

Prompt Engineering

STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents

2025-09-30 · Jing-Jing Li, Jianfeng He, Chao Shang, Devang Kulshreshtha 외 arxiv

As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chainin…

CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation

2025-08-11 · Shuzhou Yuan, William LaCroix, Hardik Ghoshal, Ercong Nie 외 arxiv

Large Language Models (LLMs) are increasingly employed as AI tutors due to their scalability and potential for personalized instruction. However, off-the-shelf LLMs often underperform in educational settings: they freque…

Data Augmentation

DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training

2026-04-18 · Ziwen Pan, Zihan Liang, Jad Kabbara, Ali Emami arxiv

Large language models (LLMs) tuned for safety often avoid acknowledging demographic differences, even when such acknowledgment is factually correct (e.g., ancestry-based disease incidence) or contextually justified (e.g.…