paper-with-me

홈 › Papers

(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable

2026-06-11 · Chen Zhu, Xiaolu Wang, Weilong Zhang arxiv

Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capability, but also on how cognitive labour is structured between humans and machines. We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation. In a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets, an unconstrained multi-agent baseline produced critical failures in 72% of runs. Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow. Fisher's exact test rejects equality of failure rates at p<0.001. Reliability gains were largest on the least publicly represented dataset, a Qing-dynasty population register, consistent with a task-based production model with Frechet-distributed output quality. An 80-run ablation suggests that deterministic computation and human gates contribute independently, with exploratory evidence of complementarity. We interpret HLER as a research harness rather than an autonomous AI scientist: it sharply reduces failures, makes residual weaknesses more visible, and prevents unreliable claims from being advanced as publication-ready outputs.

📄 PDF Abstract BibTeX arXiv:2606.12848

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight

2025-09-12 · Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang 외 arxiv

The dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents…

Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems

2026-04-09 · Susanne Gaube, Markus Langer, Tim Miller, Kevin Baum 외 arxiv

The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that may only be ameliorated by human oversight. However, notions of human ov…

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?

2026-06-04 · Jingheng Ye, Huiqi Zou, Simon Yu, Weiyan Shi arxiv

AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can expl…

Beyond Procedural Compliance: Human Oversight as a Dimension of Well-being Efficacy in AI Governance

2025-12-15 · Yao Xie, Walter Cullen arxiv

Major AI ethics guidelines and laws, including the EU AI Act, call for effective human oversight, but do not define it as a distinct and developable capacity. This paper introduces human oversight as a well-being capacit…

Nonuniformity Principle in Human-AI Coworking

2026-07-17 · An Luo, Jie Ding hf

As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvement remain essential for ensuring the quality of AI-generated outputs. In practice, while it is desirab…