paper-with-me

홈 › Papers

PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk

2026-04-13 · Seulki Lee arxiv

Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines can be set more fundamentally -- at the level of value, evidence, and source hierarchies that govern AI reasoning. Using the PRISM (Profile-based Reasoning Integrity Stack Measurement) framework, we define a taxonomy of 27 behavioral risk signals derived from structural anomalies in how AI systems prioritize values (L4), weight evidence types (L3), and trust information sources (L2). Each signal is evaluated through a dual-threshold principle combining absolute rank position and relative win-rate gap, producing a two-tier classification (Confirmed Risk vs. Watch Signal). The hierarchy-based approach offers three advantages over case-specific red lines: it is anticipatory rather than reactive (detecting dangerous reasoning structures before they produce harmful outputs), comprehensive rather than enumerative (a single value-hierarchy signal subsumes an unlimited number of case-specific violations), and measurable rather than subjective (grounded in empirical forced-choice data). We demonstrate the framework's detection capacity using approximately 397,000 forced-choice responses from 7 AI models across three Authority Stack layers, showing that the signal taxonomy successfully discriminates between models with structurally extreme profiles, models with context-dependent risk, and models with balanced hierarchies.

📄 PDF Abstract BibTeX arXiv:2604.11070

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment

2026-04-02 · Chenning Xu, Mao Zheng, Mingyang Song arxiv

Supervised fine-tuning (SFT) with token-level hard labels can amplify overconfident imitation of factually unsupported targets, causing hallucinations that propagate in multi-sentence generation. We study an augmented SF…

PRISM: A hierarchical multiscale approach for time series forecasting

2025-12-31 · Zihao Chen, Alexandre Andre, Wenrui Ma, Ian Knight 외 arxiv

Forecasting is critical in areas such as finance, biology, and healthcare. Despite the progress in the field, making accurate forecasts remains challenging because real-world time series contain both global trends, local…

Time Series Forecasting

PRISM: Generation-Time Detection and Mitigation of Secret Leakage in Multi-Agent LLM Pipelines

2026-05-11 · Riya Tapwal, Abhishek Kumar, Carsten Maple arxiv

Multi-agent LLM systems introduce a security risk in which sensitive information accessed by one agent can propagate through shared context and reappear in downstream outputs, even without explicit adversarial intent. We…

PRISM: A Framework Harnessing Unsupervised Visual Representations and Textual Prompts for Explainable MACE Survival Prediction from Cardiac Cine MRI

2025-08-26 · Haoyang Su, Jin-Yi Xiang, Shaohao Rui, Yifan Gao 외 arxiv

Accurate prediction of major adverse cardiac events (MACE) remains a central challenge in cardiovascular prognosis. We present PRISM (Prompt-guided Representation Integration for Survival Modeling), a self-supervised fra…

L-PRISMA: An Extension of PRISMA in the Era of Generative Artificial Intelligence (GenAI)

2026-01-06 · Samar Shailendra, Rajan Kadel, Aakanksha Sharma, Islam Mohammad Tahidul 외 arxiv

The Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) framework provides a rigorous foundation for evidence synthesis, yet the manual processes of data extraction and literature screening remain…