paper-with-me

Papers

Supertrust foundational alignment: mutual trust must replace permanent control for safe superintelligence

2024-07-29 · James M. Mazzu

It's widely expected that humanity will someday create AI systems vastly more intelligent than us, leading to the unsolved alignment problem of "how to control superintelligence." However, this commonly expressed problem is not only self-contradictory and likely unsolvable, but current strategies to ensure permanent control effectively guarantee that superintelligent AI will distrust humanity and consider us a threat. Such dangerous representations, already embedded in current models, will inevitably lead to an adversarial relationship and may even trigger the extinction event many fear. As AI leaders continue to "raise the alarm" about uncontrollable AI, further embedding concerns about it "getting out of our control" or "going rogue," we're unintentionally reinforcing our threat and deepening the risks we face. The rational path forward is to strategically replace intended permanent control with intrinsic mutual trust at the foundational level. The proposed Supertrust alignment meta-strategy seeks to accomplish this by modeling instinctive familial trust, representing superintelligence as the evolutionary child of human intelligence, and implementing temporary controls/constraints in the manner of effective parenting. Essentially, we're creating a superintelligent "child" that will be exponentially smarter and eventually independent of our control. We therefore have a critical choice: continue our controlling intentions and usher in a brief period of dominance followed by extreme hardship for humanity, or intentionally create the foundational mutual trust required for long-term safe coexistence.

📄 PDF Abstract BibTeX arXiv:2407.20208

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

2026-01-09 · Edward C. Cheng, Jeshua Cheng, Alice Siu arxiv

This paper presents a conceptual and operational framework for developing and operating safe and trustworthy AI agents based on a Three-Pillar Model grounded in transparency, accountability, and trustworthiness. Building…

Reinforcement LearningAutonomous Driving

Toward an Interaction-Centered Approach to Robot Trustworthiness

2025-08-19 · Carlo Mazzola, Hassan Ali, Kristína Malinovská, Igor Farkaš arxiv

As robots get more integrated into human environments, fostering trustworthiness in embodied robotic agents becomes paramount for an effective and safe human-robot interaction (HRI). To achieve that, HRI applications mus…

Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

2025-05-26 · Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Nour Aburaed 외

This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensit…

Not someone, but something: Rethinking trust in the age of medical AI

2025-04-04 · Jan Beger

As artificial intelligence (AI) becomes embedded in healthcare, trust in medical decision-making is changing fast. This opinion paper argues that trust in AI isn't a simple transfer from humans to machines - it's a dynam…

Decision MakingPhilosophy

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

2026-05-18 · Yixiang Yao, Yuhang Yao, Xinyi Fan, Jiechao Gao 외 arxiv

The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As these agents transition from isolated operation to collaborative ecosystems, we …