paper-with-me

Papers

Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard

2026-06-07 · Alireza Arbabi, Florian Kerschbaum arxiv

Large language models (LLMs) are increasingly released and deployed through opaque development and deployment pipelines, enabling model providers to inject intentional, provider-specific policies without officially announcing them. As a result, various models have been reported to generate responses reflecting proprietary rules and organizational interests, leading to censorship or misinformation on controversial topics. However, systematic identification of such alignment remains a fundamental challenge, complicated by the ambiguity of what ``proprietary'' entails in different contexts. In this paper, we propose a statistical framework for detecting proprietary alignment in black-box language models via comparative behavioral analysis. Our approach quantifies systematic deviations between the responses of a target model and those of a reference set of baseline models in a shared semantic space. By evaluating relative behavioral divergence rather than absolute correctness, our framework enables principled auditing under black-box access. Applied to several widely discussed but previously unquantified cases, it provides a systematic and scalable basis for external assessment of provider-specific alignment behavior in large language models.

📄 PDF Abstract BibTeX arXiv:2606.08381

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models

2025-07-22 · Armin Berger, Lars Hillebrand, David Leonhard, Tobias Deußer 외 arxiv

The auditing of financial documents, historically a labor-intensive process, stands on the precipice of transformation. AI-driven solutions have made inroads into streamlining this process by recommending pertinent text …

TRUST: A Decentralized Framework for Auditing Large Language Model Reasoning

2025-10-23 · Morris Yu-Chao Huang, Zhen Tan, Mohan Zhang, Pingzhi Li 외 arxiv

Large Language Models generate complex reasoning chains that reveal their decision-making, yet verifying the faithfulness and harmlessness of these intermediate steps remains a critical unsolved problem. Existing auditin…

Prompt Programming for Cultural Bias and Alignment of Large Language Models

2026-03-17 · Maksim Eren, Eric Michalak, Brian Cook, Johnny Seales arxiv

Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural biases that misalign with target populations. As LLMs are increasingly used for str…

Prompt Engineering

CoDA: Exploring Chain-of-Distribution Attacks and Post-Hoc Token-Space Repair for Medical Vision-Language Models

2026-03-19 · Xiang Chen, Fangfang Yang, Chunlei Meng, Yuxian Dong 외 arxiv

Medical vision--language models (MVLMs) are increasingly used as perceptual backbones in radiology pipelines and as the visual front end of multimodal assistants, yet their reliability under real clinical workflows remai…

Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis

2025-11-21 · Haijiang Liu, Jinguang Gu, Xun Wu, Daniel Hershcovich 외 arxiv

As Large Language Models (LLMs) increasingly influence high-stakes decision-making across global contexts, ensuring their alignment with diverse cultural values has become a critical governance challenge. This study pres…

Reinforcement Learning