paper-with-me

홈 › Papers

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

2026-07-15 · Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko, Angel Ezendu, Oluwafunke Akinbuwa, Oluwaseun Odunsi, Oluwasegun Oguntuase, Oluwadarasimi Oguntuase, Ifeoma Nwabueze, Abiodun Adereni arxiv

Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions. We introduce MamaBench, the first counterfactual benchmark for maternal and paediatric AI: 434 expert-authored clinical narratives in 217 pairs across 371 pathologies, evaluated via the Bias Trap Rate (BTR), the conditional probability that a model fails the counterfactual given success on the base case. We propose Evidence-Anchored RAG (EA-RAG), a three-stage retrieval method that replaces aggregate similarity with an evidence coverage objective through clinical parameter extraction, coverage auditing, and contrastive sub-queries. Across eight configurations of four frontier LLMs, base accuracy overstates robust accuracy by 16-28 percentage points in every model. EA-RAG achieves 20.3% BTR and 65.0% robust accuracy on Claude Sonnet 4.6, a 5.5 percentage point BTR reduction without degrading base accuracy. The residual 20% BTR confirms that counterfactual robustness in clinical AI remains an open challenge. Keywords: counterfactual evaluation, clinical AI, maternal healthcare, retrieval-augmented generation, diagnostic robustness

📄 PDF Abstract BibTeX arXiv:2607.14385

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

2026-06-28 · Yi Ren arxiv

Medical question-answering benchmarks rarely cover the maternal, neonatal, child, and reproductive-health questions a nurse-midwife asks, and, to our knowledge, no public chunk-level relevance benchmark exists for matern…

Evaluating the Financial Factors Influencing Maternal, Newborn, and Child Health in Africa

2024-02-22 · Youssef Er-Rays, Meriem M'dioud

The study investigated the impact of healthcare system efficiency on the delivery of maternal, newborn, and child services in Africa. Data Envelopment Analysis and Tobit regression were employed to assess the efficiency …

Maternal education and offspring birthweight for gestational age: the mediating effect of smoking during pregnancy

2020-04-22

Background Small for gestational age (SGA) birthweight, a risk factor of infant mortality and delayed child development, is associated with maternal educational attainment. Maternal tobacco smoking during pregnancy could…

Predicting Anemia Among Under-Five Children in Nepal Using Machine Learning and Deep Learning

2026-02-01 · Deepak Bastola, Pitambar Acharya, Dipak Dulal, Rabina Dhakal 외 arxiv

Childhood anemia remains a major public health challenge in Nepal and is associated with impaired growth, cognition, and increased morbidity. Using World Health Organization hemoglobin thresholds, we defined anemia statu…

Binary Classification

Public Health Insurance of Children and Maternal Labor Market Outcomes

2025-04-14 · Konstantin Kunze

This paper exploits variation resulting from a series of federal and state Medicaid expansions between 1977 and 2017 to estimate the effects of children's increased access to public health insurance on the labor market o…