paper-with-me

홈 › Papers

RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis

2026-01-13 · Zhengwei Tao, Bo Li, Jialong Wu, Guochen Yan, Huanyao Zhang, Jiahao Xu, Haitao Mi, Wentao Zhang arxiv

Agentic Retrieval-Augmented Generation (RAG) empowers large language models to autonomously plan and retrieve information for complex problem-solving. However, the development of robust agents is hindered by the scarcity of high-quality training data that reflects the noise and complexity of real-world retrieval environments. Conventional manual annotation is unscalable and often fails to capture the dynamic reasoning strategies required to handle retrieval failures. To bridge this gap, we introduce RAGShaper, a novel data synthesis framework designed to automate the construction of RAG tasks and robust agent trajectories. RAGShaper incorporates an InfoCurator to build dense information trees enriched with adversarial distractors spanning Perception and Cognition levels. Furthermore, we propose a constrained navigation strategy that forces a teacher agent to confront these distractors, thereby eliciting trajectories that explicitly demonstrate error correction and noise rejection. Comprehensive experiments confirm that models trained on our synthesized corpus significantly outperform existing baselines, exhibiting superior robustness in noise-intensive and complex retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2601.08699

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

2026-05-02 · Tianshi Zheng, Rui Wang, Xiyun Li, Kelvin Kiu Wai Tam 외 arxiv

Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents offer a promising approach to this challenge. These models develop rob…

Reinforcement Learning

Towards eliciting latent knowledge from LLMs with mechanistic interpretability

2025-05-20 · Bartosz Cywiński, Emil Ryd, Senthooran Rajamanoharan, Neel Nanda

As language models become more powerful and sophisticated, it is crucial that they remain trustworthy and reliable. There is concerning preliminary evidence that models may attempt to deceive or keep secrets from their o…

Skilled AI Agents for Embedded and IoT Systems Development

2026-03-20 · Yiming Li, Yuhan Cheng, Mingchen Ma, Yihang Zou 외 arxiv

Large language models (LLMs) and agentic systems have shown promise for automated software development, but applying them to hardware-in-the-loop (HIL) embedded and Internet-of-Things (IoT) systems remains challenging du…

Optimizing AI Agent Attacks With Synthetic Data

2025-11-04 · Chloe Loughridge, Paul Colognese, Avery Griffin, Tyler Tracy 외 arxiv

As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good control evaluations require eliciting st…

Agentic Uncertainty Reveals Agentic Overconfidence

2026-02-06 · Jean Kaddour, Srijan Patel, Gbètondji Dovonon, Leo Richter 외 arxiv

Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All results exhibit agentic overconfidence: som…