paper-with-me

Papers

BenchmarkCards: Large Language Model and Risk Reporting

2024-10-16 · Anna Sokol, Nuno Moniz, Elizabeth Daly, Michael Hind, Nitesh Chawla

Large language models (LLMs) offer powerful capabilities but also introduce significant risks. One way to mitigate these risks is through comprehensive pre-deployment evaluations using benchmarks designed to test for specific vulnerabilities. However, the rapidly expanding body of LLM benchmark literature lacks a standardized method for documenting crucial benchmark details, hindering consistent use and informed selection. BenchmarkCards addresses this gap by providing a structured framework specifically for documenting LLM benchmark properties rather than defining the entire evaluation process itself. BenchmarkCards do not prescribe how to measure or interpret benchmark results (e.g., defining ``correctness'') but instead offer a standardized way to capture and report critical characteristics like targeted risks and evaluation methodologies, including properties such as bias and fairness. This structured metadata facilitates informed benchmark selection, enabling researchers to choose appropriate benchmarks and promoting transparency and reproducibility in LLM evaluation.

📄 PDF Abstract BibTeX arXiv:2410.12974

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessLanguage ModelingLanguage ModellingLarge Language Modelmodel

Similar Papers 제목 키워드 기반

Supervision policies can shape long-term risk management in general-purpose AI models

2025-01-10 · Manuel Cebrian, Emilia Gomez, David Fernandez Llorca

The rapid proliferation and deployment of General-Purpose AI (GPAI) models, including large language models (LLMs), present unprecedented challenges for AI supervisory entities. We hypothesize that these entities will ne…

DiversityManagementNavigate

MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models

2025-06-09 · Philip R. Liu, Sparsh Bansal, Jimmy Dinh, Aditya Pawar 외

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applyin…

DiagnosticHallucination

Responsible Reporting for Frontier AI Development

2024-04-03 · Noam Kolt, Markus Anderljung, Joslyn Barnhart, Asher Brass 외

Mitigating the risks from frontier AI systems requires up-to-date and reliable information about those systems. Organizations that develop and deploy frontier systems have significant access to such information. By repor…

Management

Learning Survival Models with Right-Censored Reporting Delays

2025-10-06 · Yuta Shikuri, Hironori Fujisawa arxiv

Survival analysis provides statistical methods to model the time until an event occurs. Reporting delays arise when event times are not observed at their occurrence but are only revealed upon reporting. This issue is par…

Profit and loss decomposition in continuous time and approximations

2022-12-13 · Gero Junike, Hauke Stier, Marcus C. Christiansen

Financial institutions and insurance companies that analyze the evolution and sources of profits and losses often look at risk factors only at discrete reporting dates, ignoring the detailed paths. Continuous-time decomp…