paper-with-me

홈 › Papers

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

2026-03-24 · Abhishek Chandwani, Ishan Gupta arxiv

Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world enterprise work is often subjective and context-dependent: success hinges on organizational goals, user intent, and the quality of intermediate artifacts produced across long, multi-tool workflows. We introduce LH-Bench, a three-pillar evaluation design that moves beyond binary correctness to score autonomous, long-horizon execution on subjective enterprise tasks. The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation. We show that domain-authored rubrics provide substantially more reliable evaluation signals than LLM-authored rubrics (kappa = 0.60 vs. 0.46), and that human preference judgments confirm the same top-tier separation (p < 0.05), evidence that expert-grounded evaluation can scale without sacrificing reliability. We release public datasets and report results on two environments: Figma-to-code (33 real .fig tasks against the Figma API via MCP) and Programmatic content (41 courses comprising 183 individually-evaluated chapters on a course platform serving 30+ daily users).

📄 PDF Abstract BibTeX arXiv:2603.22744

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks

2026-08-31 · Chunyun Ma, Lun Luo, Xingjian Luo, Xiexing Feng 외 arxiv

Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success depends on the successful completion of multiple constituent skills. Existing benchmarks, however, still rely …

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

2026-08-05 · Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang 외 hf

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems…

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

2026-08-06 · Zhi Han, Chenxi Zeng, Liuhaichen Yang, Zihan Guo 외 arxiv

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verific…

Spatially Grounded Long-Horizon Task Planning in the Wild

2026-03-13 · Sehun Jung, HyunJee Song, Dong-Hee Kim, Reuben Tan 외 arxiv

Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action plans expressed in natural language that gu…

Robot Manipulation

Lifting Traces to Logic: Programmatic Skill Induction with Neuro-Symbolic Learning for Long-Horizon Agentic Tasks

2026-05-02 · Jie-Jing Shao, Haiyan Yin, Yueming Lyu, Xingrui Yu 외 arxiv

Foundation model-driven agents often struggle with long-horizon planning due to the transient nature of purely prompting-based reasoning. While existing skill induction methods mitigate this by distilling experience into…