paper-with-me

Papers

Agents on the Bench: Large Language Model Based Multi Agent Framework for Trustworthy Digital Justice

2024-12-24 · Cong Jiang, Xiaolei Yang

The justice system has increasingly employed AI techniques to enhance efficiency, yet limitations remain in improving the quality of decision-making, particularly regarding transparency and explainability needed to uphold public trust in legal AI. To address these challenges, we propose a large language model based multi-agent framework named AgentsBench, which aims to simultaneously improve both efficiency and quality in judicial decision-making. Our approach leverages multiple LLM-driven agents that simulate the collaborative deliberation and decision making process of a judicial bench. We conducted experiments on legal judgment prediction task, and the results show that our framework outperforms existing LLM based methods in terms of performance and decision quality. By incorporating these elements, our framework reflects real-world judicial processes more closely, enhancing accuracy, fairness, and society consideration. AgentsBench provides a more nuanced and realistic methods of trustworthy AI decision-making, with strong potential for application across various case types and legal scenarios.

📄 PDF Abstract BibTeX arXiv:2412.18697

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingFairnessLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Uphold 설명 없음

Similar Papers 제목 키워드 기반

X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System

2025-05-21 · Peng Wang, Ruihan Tao, Qiguang Chen, Mengkang Hu 외

Recently, large language model (LLM)-based agents have achieved significant success in interactive environments, attracting significant academic and industrial attention. Despite these advancements, current research pred…

Language ModelingLanguage ModellingLarge Language Model

Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level

2024-04-08 · Chenxu Wang, Bin Dai, Huaping Liu, Baoyuan Wang

Prominent large language models have exhibited human-level performance in many domains, even enabling the derived agents to simulate human and social interactions. While practical works have substantiated the practicabil…

Benchmarking

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

2025-02-13 · Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao 외

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-bas…

Benchmarking

Benchmarking Mobile Device Control Agents across Diverse Configurations

2024-04-25 · Juyong Lee, Taywon Min, Minyong An, Dongyoon Hahm 외

Mobile device control agents can largely enhance user interactions and productivity by automating daily tasks. However, despite growing interest in developing practical agents, the absence of a commonly adopted benchmark…

BenchmarkingImitation Learning

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

2026-05-09 · Tao Yu, Hao Wang, Changyu Li, Shenghua Chai 외 arxiv

Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. How…