paper-with-me

홈 › Papers

The AI Alignment Paradox

2024-05-31 · Robert West, Roland Aydin

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This perspective article draws attention to a fundamental challenge we see in all AI alignment endeavors, which we term the "AI alignment paradox": The better we align AI models with our values, the easier we may make it for adversaries to misalign the models. We illustrate the paradox by sketching three concrete example incarnations for the case of language models, each corresponding to a distinct way in which adversaries might exploit the paradox. With AI's increasing real-world impact, it is imperative that a broad community of researchers be aware of the AI alignment paradox and work to find ways to mitigate it, in order to ensure the beneficial use of AI for the good of humanity.

📄 PDF Abstract BibTeX arXiv:2405.20806

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Towards Adaptive Unknown Authentication for Universal Domain Adaptation by Classifier Paradox

2022-07-10 · Yunyun Wang, Yao Liu, Songcan Chen

Universal domain adaptation (UniDA) is a general unsupervised domain adaptation setting, which addresses both domain and label shifts in adaptation. Its main challenge lies in how to identify target samples in unshared o…

Domain AdaptationUniversal Domain AdaptationUnsupervised Domain Adaptation

Persona-aware Generative Model for Code-mixed Language

2023-09-06 · Ayan Sengupta, Md Shad Akhtar, Tanmoy Chakraborty

Code-mixing and script-mixing are prevalent across online social networks and multilingual societies. However, a user's preference toward code-mixing depends on the socioeconomic status, demographics of the user, and the…

Decodermodelvalid

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

2026-05-23 · Dingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller 외 arxiv

While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invariants, making even perception-based out-of…

Representation Learning

The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models

2024-10-09 · Yanjun Chen, Dawei Zhu, Yirong Sun, Xinghao Chen 외

Reinforcement Learning from Human Feedback significantly enhances Natural Language Processing by aligning language models with human expectations. A critical factor in this alignment is the strength of reward models used…

The Democratic Paradox in Large Language Models' Underestimation of Press Freedom

2025-06-22 · I. Loaiza, R. Vestrelli, A. Fronzetti Colladon, R. Rigobon

As Large Language Models (LLMs) increasingly mediate global information access for millions of users worldwide, their alignment and biases have the potential to shape public understanding and trust in fundamental democra…