paper-with-me

홈 › Papers

RAVine: Reality-Aligned Evaluation for Agentic Search

2025-07-22 · Yilong Xu, Xiang Long, Zhi Zheng, Jinhua Gao arxiv

Agentic search, as a more autonomous and adaptive paradigm of retrieval augmentation, is driving the evolution of intelligent search systems. However, existing evaluation frameworks fail to align well with the goals of agentic search. First, the complex queries commonly used in current benchmarks often deviate from realistic user search scenarios. Second, prior approaches tend to introduce noise when extracting ground truth for end-to-end evaluations, leading to distorted assessments at a fine-grained level. Third, most current frameworks focus solely on the quality of final answers, neglecting the evaluation of the iterative process inherent to agentic search. To address these limitations, we propose RAVine -- a Reality-Aligned eValuation framework for agentic LLMs with search. RAVine targets multi-point queries and long-form answers that better reflect user intents, and introduces an attributable ground truth construction strategy to enhance the accuracy of fine-grained evaluation. Moreover, RAVine examines model's interaction with search tools throughout the iterative process, and accounts for factors of efficiency. We benchmark a series of models using RAVine and derive several insights, which we hope will contribute to advancing the development of agentic search systems. The code and datasets are available at https://github.com/SwordFaith/RAVine.

📄 PDF Abstract BibTeX arXiv:2507.16725

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Predator-Prey Model: Driven Hunt for Accelerated Grokking

2025-09-10 · I. A. Lopatin, S. V. Kozyrev, A. N. Pechen arxiv

A machine learning method is proposed using two agents that simulate the biological behavior of a predator and a prey. In this method, the predator and the prey interact with each other - the predator chases the prey whi…

UrbanWorld2.0: A Multimodal Agentic Framework for Reality-Aligned 3D World Generation at City-Scale

2025-11-22 · Shengyuan Wang, Zhiheng Zheng, Yu Shang, Lixuan He 외 arxiv

The automated generation of high-fidelity, city-scale 3D environments remains a formidable challenge with profound academic and industrial implications. However, existing methods struggle to achieve the necessary quality…

3D Generation

Ravines in quantum cost landscapes: opportunities for improved VQA predictions

2026-07-01 · Felix J. Beckmann, João F. Bravo arxiv

The geometric and topological structure of quantum cost landscapes (QCLs) governs the optimization and thus the predictive power of variational quantum algorithms (VQAs). We systematically analyze ravines - low-cost path…

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

2026-04-04 · Erhan Zhang, Yiqun Chen, Zechun Niu, Wei Yang 외 arxiv

Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforcement learning with verifiable rewards (RLVR) has emerged as a widely adopt…

Reinforcement Learning

LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation

2026-02-18 · Hejia Zhang, Zhongming Yu, Chia-Tung Ho, Haoxing Ren 외 arxiv

Execution-aware LLM agents offer a promising paradigm for learning from tool feedback, but such feedback can be expensive and slow to obtain, making online reinforcement learning (RL) less practical in certain scenarios.…

Reinforcement Learning