paper-with-me

홈 › Papers

Advancing ESG Intelligence: An Expert-level Agent and Comprehensive Benchmark for Sustainable Finance

2026-01-13 · Yilei Zhao, Wentao Zhang, Lei Xiao, Yandan Zheng, Mengpu Liu, Wei Yang Bryan Lim arxiv

Environmental, social, and governance (ESG) criteria are essential for evaluating corporate sustainability and ethical performance. However, professional ESG analysis is hindered by data fragmentation across unstructured sources, and existing large language models (LLMs) often struggle with the complex, multi-step workflows required for rigorous auditing. To address these limitations, we introduce ESGAgent, a hierarchical multi-agent system empowered by a specialized toolset, including retrieval augmentation, web search and domain-specific functions, to generate in-depth ESG analysis. Complementing this agentic system, we present a comprehensive three-level benchmark derived from 310 corporate sustainability reports, designed to evaluate capabilities ranging from atomic common-sense questions to the generation of integrated, in-depth analysis. Empirical evaluations demonstrate that ESGAgent outperforms state-of-the-art closed-source LLMs with an average accuracy of 84.15% on atomic question-answering tasks, and excels in professional report generation by integrating rich charts and verifiable references. These findings confirm the diagnostic value of our benchmark, establishing it as a vital testbed for assessing general and advanced agentic capabilities in high-stakes vertical domains.

📄 PDF Abstract BibTeX arXiv:2601.08676

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments

2025-07-05 · Zheng Jia, Shengbin Yue, Wei Chen, Siyuan Wang 외 arxiv

The gap between static benchmarks and the dynamic nature of real-world legal practice poses a key barrier to advancing legal intelligence. To this end, we introduce J1-ENVS, the first interactive and dynamic legal enviro…

Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles

2026-02-02 · Shaohan Wang, Benfeng Xu, Licheng Zhang, Mingxuan Du 외 arxiv

Deep Research Agents (DRAs) have demonstrated remarkable capabilities in autonomous information retrieval and report generation, showing great potential to assist humans in complex research tasks. Current evaluation fram…

Information Retrieval

HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-Making

2025-09-16 · Xingxing Hong, Yungong Wang, Dexin Jin, Ye Yuan 외 arxiv

Benchmarks are crucial for assessing multi-agent reinforcement learning (MARL) algorithms. While StarCraft II-related environments have driven significant advances in MARL, existing benchmarks like SMAC focus primarily o…

Multi-agent Reinforcement LearningStarcraft II

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

2026-05-31 · Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo 외 arxiv

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are i…

Enhanced Classroom Dialogue Sequences Analysis with a Hybrid AI Agent: Merging Expert Rule-Base with Large Language Models

2024-11-13 · Yun Long, Yu Zhang

Classroom dialogue plays a crucial role in fostering student engagement and deeper learning. However, analysing dialogue sequences has traditionally relied on either theoretical frameworks or empirical descriptions of pr…

AI AgentLanguage ModelingLanguage ModellingLarge Language Model