paper-with-me

홈 › Papers

Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark

2024-11-28 · Jianyou Wang, Weili Cao, Longtian Bao, Youze Zheng, Gil Pasternak, Kaicheng Wang, Xiaoyue Wang, Ramamohan Paturi, Leon Bergen

Systems that answer questions by reviewing the scientific literature are becoming increasingly feasible. To draw reliable conclusions, these systems should take into account the quality of available evidence, placing more weight on studies that use a valid methodology. We present a benchmark for measuring the methodological strength of biomedical papers, drawing on the risk-of-bias framework used for systematic reviews. The four benchmark tasks, drawn from more than 500 papers, cover the analysis of research study methodology, followed by evaluation of risk of bias in these studies. The benchmark contains 2000 expert-generated bias annotations, and a human-validated pipeline for fine-grained alignment with research paper content. We evaluate a range of large language models on the benchmark, and find that these models fall significantly short of expert-level performance. By providing a standardized tool for measuring judgments of study quality, the benchmark can help to guide systems that perform large-scale aggregation of scientific data. The dataset is available at https://github.com/RoBBR-Benchmark/RoBBR.

📄 PDF Abstract BibTeX arXiv:2411.18831

Code (1)

robbr-benchmark/robbr 공식 구현 pytorch

Tasks

valid

Similar Papers 제목 키워드 기반

Quantifying 60 Years of Gender Bias in Biomedical Research with Word Embeddings

2020-07-01 · WS 2020 7 · Anthony Rios, Reenam Joshi, Hejin Shin

Gender bias in biomedical research can have an adverse impact on the health of real people. For example, there is evidence that heart disease-related funded research generally focuses on men. Health disparities can form …

ArticlesWord Embeddings

Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models

2023-05-11 · Lukáš Mikula, Michal Štefánik, Marek Petrovič, Petr Sojka

While the Large Language Models (LLMs) dominate a majority of language understanding tasks, previous work shows that some of these results are supported by modelling spurious correlations of training datasets. Authors co…

Question Answering

Addressing Exposure Bias With Document Minimum Risk Training: Cambridge at the WMT20 Biomedical Translation Task

2020-10-11 · WMT (EMNLP) 2020 11 · Danielle Saunders, Bill Byrne

The 2020 WMT Biomedical translation task evaluated Medline abstract translations. This is a small-domain translation task, meaning limited relevant training data with very distinct style and vocabulary. Models trained on…

SentenceTranslation

BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery

2026-06-19 · Jieyi Wang, Bingxuan Li, Nanyi Jiang, Desong Meng 외 hf

Biomedical researchers increasingly use AI-generated analyses and reports to interpret protein-level signals, but static outputs are often insufficient for research decision-making, where users need to inspect evidence, …

Measuring Gender Bias in Word Embeddings across Domains and Discovering New Gender Bias Word Categories

2019-08-01 · WS 2019 8 · Kaytlin Chaloner, Alfredo Maldonado

Prior work has shown that word embeddings capture human stereotypes, including gender bias. However, there is a lack of studies testing the presence of specific gender bias categories in word embeddings across diverse do…

Bias DetectionClusteringGender Bias DetectionTwo-sample testing+1