paper-with-me

Papers

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

2026-05-26 · Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu arxiv

Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and largely remain to text-only interaction, making it difficult to tell whether an agent's language is actually grounded in what it perceived and did, or to identify the failure modes underlying its behavior. To address this gap, we introduce QUACK, an open-source environment and evaluation framework for auditing the grounding of agent language in multimodal social reasoning. QUACK evaluates agents at three levels: game outcomes, behavioral trajectories, and utterance-level consistency. Its core Statement Verification Pipeline reconstructs each agent's ground-truth trajectory from engine logs and checks every discussion claim against it, automatically flagging spatial hallucination, unsupported accusation, deception collapse, and language-action inconsistency. Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, we find that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and 11.5% of accusations are strictly unsupported. We release the full engine, evaluation framework, toolkit, and logs in https://github.com/AAAAA-Academia-Attractions/QUACK.

📄 PDF Abstract BibTeX arXiv:2605.27068

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content

2026-06-02 · Yang Zhao, Yingshuo Li, Zeyu Zhang arxiv

Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation. In these settings, users may obtain highly structured answers, plans, and judgments wi…

Code Generation

Bound by the Bounty: Collaboratively Shaping Evaluation Processes for Queer AI Harms

2023-07-15 · Organizers Of QueerInAI, Nathan Dennler, Anaelia Ovalle, Ashwin Singh 외

Bias evaluation benchmarks and dataset and model documentation have emerged as central processes for assessing the biases and harms of artificial intelligence (AI) systems. However, these auditing processes have been cri…

QuACK: Accelerating Gradient-Based Quantum Optimization with Koopman Operator Learning

2022-11-02 · NeurIPS 2023 11 · Di Luo, Jiayu Shen, Rumen Dangovski, Marin Soljačić

Quantum optimization, a key application of quantum computing, has traditionally been stymied by the linearly increasing complexity of gradient calculations with an increasing number of parameters. This work bridges the g…

Operator learningQuantum Machine Learning

Introspective Growth: Automatically Advancing LLM Expertise in Technology Judgment

2025-05-18 · Siyang Wu, Honglin Bao, Nadav Kunievsky, James A. Evans

Large language models (LLMs) increasingly demonstrate signs of conceptual understanding, yet much of their internal knowledge remains latent, loosely structured, and difficult to access or evaluate. We propose self-quest…

Diagnostic

QUACK: Quantum Aligned Centroid Kernel

2024-05-01 · Kilian Tscharke, Sebastian Issel, Pascal Debus

Quantum computing (QC) seems to show potential for application in machine learning (ML). In particular quantum kernel methods (QKM) exhibit promising properties for use in supervised ML tasks. However, a major disadvanta…

Dimensionality Reduction