paper-with-me

홈 › Papers

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

2026-02-06 · Lanbo Lin, Jiayao Liu, Tianyuan Yang, Li Cai, Yuanwu Xu, Lei Wei, Sicong Xie, Guannan Zhang arxiv

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response strategies, while LLM-as-a-judge approaches adapt to individual responses yet suffer from instability and bias. Human experts address this dilemma by combining domain-grounded principles with dynamic, claim-level assessment. Inspired by this process, we propose JADE, a two-layer evaluation framework. Layer 1 encodes expert knowledge as a predefined set of evaluation skills, providing stable evaluation criteria. Layer 2 performs report-specific, claim-level evaluation to flexibly assess diverse reasoning strategies, with evidence-dependency gating to invalidate conclusions built on refuted claims. Experiments on BizBench show that JADE improves evaluation stability and reveals critical agent failure modes missed by holistic LLM-based evaluators. We further demonstrate strong alignment with expert-authored rubrics and effective transfer to HealthBench and DR.BENCH, covering medical and 10-domain professional evaluation settings. Code and data are available at https://github.com/smiling-world/JADE.

📄 PDF Abstract BibTeX arXiv:2602.06486

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models

2023-11-01 · Mi Zhang, Xudong Pan, Min Yang

In this paper, we present JADE, a targeted linguistic fuzzing platform which strengthens the linguistic complexity of seed questions to simultaneously and consistently break a wide range of widely-used LLMs categorized i…

Natural Questions

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring

2025-08-28 · Junjie Chu, Mingjie Li, Ziqing Yang, Ye Leng 외 arxiv

Accurately determining whether a jailbreak attempt has succeeded is a fundamental yet unresolved challenge. Existing evaluation methods rely on misaligned proxy indicators or naive holistic judgments. They frequently mis…

JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG

2026-01-29 · Yiqun Chen, Erhan Zhang, Tianyi Hu, Shijie Wang 외 arxiv

The evolution of Retrieval-Augmented Generation (RAG) has shifted from static retrieval pipelines to dynamic, agentic workflows where a central planner orchestrates multi-turn reasoning. However, existing paradigms face …

A Collaborative Jade Recognition System for Mobile Devices Based on Lightweight and Large Models

2025-02-20 · Zhenyu Wang, Wenjia Li, Pengyu Zhu

With the widespread adoption and development of mobile devices, vision-based recognition applications have become a hot topic in research. Jade, as an important cultural heritage and artistic item, has significant applic…

JADE: Corpus for Japanese Definition Modelling

2022-06-01 · LREC 2022 6 · Han Huang, Tomoyuki Kajiwara, Yuki Arase

This study investigated and released the JADE, a corpus for Japanese definition modelling, which is a technique that automatically generates definitions of a given target word and phrase. It is a crucial technique for pr…

Definition Modelling