paper-with-me

Papers

From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents

2026-03-26 · Shuoling Liu, Zhiquan Tan, Kun Yi, Hui Wu, Yihan Li, Jiangpeng Yan, Liyuan Chen, Kai Chen, Qiang Yang arxiv

Deep Research Agents (DRAs) aim to answer complex questions by searching the web, checking evidence, and synthesizing conclusions across heterogeneous sources. We introduce a category-theoretic framework for evaluating and improving such agents. The framework treats deep research as a structured mapping from user intent to evidence-grounded conclusions, making retrieval traces, cross-source alignment, and final synthesis explicit. Guided by this view, we derive a mechanism-aware benchmark of 296 bilingual questions. The benchmark targets four structural skills central to real research: following multi-hop evidence chains, verifying claims across sources, re-ordering fragmented information, and rejecting unsupported assumptions. We evaluate 16 frontier systems with human verification and find that these structural tasks remain highly challenging: the best system reaches only 19.9% average accuracy. The results show that strong agents can sometimes reorganize evidence and detect false premises, but still struggle with long-horizon synthesis and intersection-heavy verification. Beyond evaluation, the same theory also leads to practical system improvements. We instantiate theory-guided interventions such as tracked search, which preserves retrieval traces, and category tools, which add explicit verification and synthesis steps. These interventions yield measurable gains in API-based deep research systems. Our work therefore provides both a challenging benchmark and concrete design guidance for building more reliable research agents.

📄 PDF Abstract BibTeX arXiv:2603.25342

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation

2026-05-14 · GAng Peng arxiv

Holistic evaluation scores capture overall output quality but do not distinguish whether a model reproduced the structural form of a user's request from whether it preserved the user's specific intent. We propose a dimen…

Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction

2026-04-24 · Yongxiang Lian, Yueyang Cang, Pingge Hu, Yuchen He 외 arxiv

Objective: Decoding visual information from electroencephalography (EEG) is an important problem in neuroscience and brain-computer interface (BCI) research. Existing methods are largely restricted to natural images and …

Contrastive LearningImage Generation

Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock

2026-01-13 · Didier Sornette, Sandro Claudio Lera, Ke Wu arxiv

Recent reports of large language models (LLMs) exhibiting behaviors such as deception, threats, or blackmail are often interpreted as evidence of alignment failure or emergent malign agency. We argue that this interpreta…

Intent Signal Theory: A Computational Framework for Intent-State Control in Human-AI Interaction

2026-05-24 · Gang Peng arxiv

Current AI interaction models treat the prompt as the primary object of exchange, omitting a critical layer: the user's latent source intent, the goal state preceding and motivating the prompt. Here we introduce Intent S…

Prompt Engineering

Evidence Transfer for Improving Clustering Tasks Using External Categorical Evidence

2018-11-09 · Athanasios Davvetas, Iraklis A. Klampanos, Vangelis Karkaletsis

In this paper we introduce evidence transfer for clustering, a deep learning method that can incrementally manipulate the latent representations of an autoencoder, according to external categorical evidence, in order to …

ClusteringRepresentation Learning