paper-with-me

홈 › Papers

TrialCalibre: A Fully Automated Causal Engine for RCT Benchmarking and Observational Trial Calibration

2026-04-28 · Amir Habibdoust, Xing Song arxiv

Real-world evidence (RWE) studies that emulate target trials increasingly inform regulatory and clinical decisions, yet residual, hard-to-quantify biases still limit their credibility. The recently proposed BenchExCal framework addresses this challenge via a two-stage Benchmark, Expand, Calibrate process, which first compares an observational emulation against an existing randomized controlled trial (RCT), then uses observed divergence to calibrate a second emulation for a new indication causal effect estimation. While methodologically powerful, BenchExCal is resource intensive and difficult to scale. We introduce TrialCalibre, a conceptualized multiagent system designed to automate and scale the BenchExCal workflow. Our framework features specialized agents such as the Orchestrator, Protocol Design, Data Synthesis, Clinical Validation, and Quantitative Calibration Agents that coordi-nate the the overall process. TrialCalibre incorpo-rates agent learning (e.g., RLHF) and knowledge blackboards to support adaptive, auditable, and transparent causal effect estimation.

📄 PDF Abstract BibTeX arXiv:2604.25832

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Protein Engineering Tournament: An Open Science Benchmark for Protein Modeling and Design

2023-09-18 · Chase Armer, Hassan Kane, Dana Cortade, Dave Estell 외

The grand challenge of protein engineering is the development of computational models that can characterize and generate protein sequences for any arbitrary function. However, progress today is limited by lack of 1) benc…

Benchmarking

SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks

2025-06-13 · Hwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming Zhang

Rigorous security-focused evaluation of large language model (LLM) agents is imperative for establishing trust in their safe deployment throughout the software development lifecycle. However, existing benchmarks largely …

BenchmarkingLarge Language Model

Causally-Guided Automated Feature Engineering with Multi-Agent Reinforcement Learning

2026-02-18 · Arun Vignesh Malarkkan, Wangyang Ying, Yanjie Fu arxiv

Automated feature engineering (AFE) enables AI systems to autonomously construct high-utility representations from raw tabular data. However, existing AFE methods rely on statistical heuristics, yielding brittle features…

Multi-agent Reinforcement LearningFeature Engineering

RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms

2026-03-05 · Kenan Li, Rongzhi Li, Linghao Zhang, Qirui Jin 외 arxiv

Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In thi…

Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities

2026-04-26 · Yichen Xu arxiv

Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference. We sho…

Causal Inference