paper-with-me

Papers

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

2026-09-10 · Nitesh V. Chawla, Paulo Benanti arxiv

Claims that Artificial Intelligence systems improve decisions, broaden access, reduce harm, or empower users can exceed what their evaluation establishes. Predictive performance alone does not establish safety, the presence of oversight does not establish meaningful control, and faster task completion does not establish understanding or choice. Evaluation must account for unreliable outputs and uneven performance, but also for overreliance, weakened recourse, and displaced human expertise. The harder questions are what the evidence warrants, which relations of power remain unexamined, and where measurement must stop. Assessing improvement requires examining what institutions value and the conditions AI is asked to address. AI is both revelation and intervention. Its use can reveal unmet human needs and assumptions about what matters. Once deployed, it can repair, compound, substitute for, or conceal existing failures. We develop a rupture test that evaluates deployment against explicit human and non-AI baselines. Drawing on Pope Leo XIV's Magnifica Humanitas, we examine dignity and the common good alongside questions of who owns AI infrastructure and who controls its use. These commitments shape judgments about improvement; evidence alone cannot establish moral or political legitimacy. We distinguish evidence-bounded deployment, which limits claims to what has been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. RISE AI provides an evidence architecture for making bounded claims about Responsibility, Inclusivity, Safety, and Empowerment. It records what is claimed, who answers for it, what evidence supports it, and what would require the claim to be qualified, revised, or withdrawn.

📄 PDF Abstract BibTeX arXiv:2609.11910

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evidence-Bound Autonomous Research (EviBound): A Governance Framework for Eliminating False Claims

2025-10-28 · Ruiying Chen arxiv

LLM-based autonomous research agents report false claims: tasks marked "complete" despite missing artifacts, contradictory metrics, or failed executions. EviBound is an evidence-bound execution framework that eliminates …

Look Again Before You Abstain:Budgeted Conformal Evidence Acquisition for Reliable Vision-Language Model

2026-06-15 · Jian Xu, Delu Zeng, John Paisley, Qibin Zhao arxiv

Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective prediction with a distribution-free guarantee-verify each claim and abstain w…

Testing and Evaluation of Agentic AI Systems In Military Command and Control

2026-08-20 · Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke 외 arxiv

Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assuran…

RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

2025-12-28 · Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli arxiv

Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce RW-Post, a post-aligned text--image benchmark for real-world multimodal …

Visual Grounding

Mental Health AI Safety Claims Must Preserve Temporal Evidence

2026-05-09 · Srimonti Dutta, Ratna Kandala arxiv

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failure…