paper-with-me

홈 › Papers

Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability

2024-12-24 · Haonan Li, Xudong Han, Zenan Zhai, Honglin Mu, Hao Wang, Zhenxuan Zhang, Yilin Geng, Shom Lin, Renxi Wang, Artem Shelmanov, Xiangyu Qi, Yuxia Wang, Donghai Hong, Youliang Yuan, Meng Chen, Haoqin Tu, Fajri Koto, Tatsuki Kuribayashi, Cong Zeng, Rishabh Bhardwaj, Bingchen Zhao, Yawen Duan, Yi Liu, Emad A. Alghamdi, Yaodong Yang, Yinpeng Dong, Soujanya Poria, PengFei Liu, Zhengzhong Liu, Xuguang Ren, Eduard Hovy, Iryna Gurevych, Preslav Nakov, Monojit Choudhury, Timothy Baldwin

To address this gap, we introduce Libra-Leaderboard, a comprehensive framework designed to rank LLMs through a balanced evaluation of performance and safety. Combining a dynamic leaderboard with an interactive LLM arena, Libra-Leaderboard encourages the joint optimization of capability and safety. Unlike traditional approaches that average performance and safety metrics, Libra-Leaderboard uses a distance-to-optimal-score method to calculate the overall rankings. This approach incentivizes models to achieve a balance rather than excelling in one dimension at the expense of some other ones. In the first release, Libra-Leaderboard evaluates 26 mainstream LLMs from 14 leading organizations, identifying critical safety challenges even in state-of-the-art models.

📄 PDF Abstract BibTeX arXiv:2412.18551

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards

2024-07-04 · Zhimin Zhao, Abdul Ali Bangash, Filipe Roseiro Côgo, Bram Adams 외

Foundation models (FM), such as large language models (LLMs), which are large-scale machine learning (ML) models, have demonstrated remarkable adaptability in various downstream software engineering (SE) tasks, such as c…

Code Completion

Automated Mining of Leaderboards for Empirical AI Research

2021-08-31 · Salomon Kabongo, Jennifer D'Souza, Sören Auer

With the rapid growth of research publications, empowering scientists to keep oversight over the scientific progress is of paramount importance. In this regard, the Leaderboards facet of information organization provides…

Knowledge GraphsScientific Results Extraction

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

2025-12-11 · Aileen Cheng, Alon Jacovi, Amir Globerson, Ben Golan 외 arxiv

We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually accurate text across diverse scenarios.…

ClonEval: An Open Voice Cloning Benchmark

2025-04-29 · Iwona Christop, Tomasz Kuczyński, Marek Kubis

We present a novel benchmark for voice cloning text-to-speech models. The benchmark consists of an evaluation protocol, an open-source library for assessing the performance of voice cloning models, and an accompanying le…

text-to-speechText to SpeechVoice Cloning

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

2025-08-28 · Sam Jung, Agustin Garcinuno, Spencer Mateega arxiv

AI text-to-app tools promise high quality applications and websites in minutes, yet no public benchmark rigorously verifies those claims. We introduce UI-Bench, the first large-scale benchmark that evaluates visual excel…