paper-with-me

홈 › Papers

Unlocking the Power of Multi-institutional Data: Integrating and Harmonizing Genomic Data Across Institutions

2024-01-30 · Yuan Chen, Ronglai Shen, Xiwen Feng, Katherine Panageas

Cancer is a complex disease driven by genomic alterations, and tumor sequencing is becoming a mainstay of clinical care for cancer patients. The emergence of multi-institution sequencing data presents a powerful resource for learning real-world evidence to enhance precision oncology. GENIE BPC, led by the American Association for Cancer Research, establishes a unique database linking genomic data with clinical information for patients treated at multiple cancer centers. However, leveraging such multi-institutional sequencing data presents significant challenges. Variations in gene panels result in loss of information when the analysis is conducted on common gene sets. Additionally, differences in sequencing techniques and patient heterogeneity across institutions add complexity. High data dimensionality, sparse gene mutation patterns, and weak signals at the individual gene level further complicate matters. Motivated by these real-world challenges, we introduce the Bridge model. It uses a quantile-matched latent variable approach to derive integrated features to preserve information beyond common genes and maximize the utilization of all available data while leveraging information sharing to enhance both learning efficiency and the model's capacity to generalize. By extracting harmonized and noise-reduced lower-dimensional latent variables, the true mutation pattern unique to each individual is captured. We assess the model's performance and parameter estimation through extensive simulation studies. The extracted latent features from the Bridge model consistently excel in predicting patient survival across six cancer types in GENIE BPC data.

📄 PDF Abstract BibTeX arXiv:2402.00077

Code (1)

ychen178/Bridge_model 공식 구현 pytorch

Tasks

parameter estimation

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning

2026-05-13 · Daniel M. Jimenez-Gutierrez, Enrique Zuazua, Georgios Kellaris, Joaquin del Rio 외 arxiv

The recent success of large language models (LLMs) has been largely driven by vast public datasets. However, the next frontier for LLM development lies beyond public data. Much of the world's most valuable information is…

parameter-efficient fine-tuningFederated LearningQuestion Answering

Levers of Power in the Field of AI

2025-11-05 · Tammy Mackenzie, Sukriti Punj, Natalie Perez, Sreyoshi Bhaduri 외 arxiv

This paper examines how decision makers in academia, government, business, and civil society navigate questions of power in implementations of artificial intelligence. The study explores how individuals experience and ex…

Dino-NestedUNet: Unlocking Foundation Vision Encoders for Pathology Tumor Bulk Segmentation via Dense Decoding

2026-04-27 · Tianyang Wang, Ziyu Su, Abdul Rehman Akbar, Usama Sajjad 외 arxiv

Vision foundation models (VFMs), such as DINOv3, provide rich semantic representations that are promising for computational pathology. However, many current adaptations pair frozen VFMs with lightweight decoders, creatin…

Health+: Empowering Individuals via Unifying Health Data

2026-02-22 · Sujaya Maiyya, Shantanu Sharma, Avinash Kumar arxiv

Managing personal health data is a challenge in today's fragmented and institution-centric healthcare ecosystem. Individuals often lack meaningful control over their medical records, which are scattered across incompatib…

Strategic priorities for transformative progress in advancing biology with proteomics and artificial intelligence

2025-02-21 · Yingying Sun, Jun A, Zhiwei Liu, Rui Sun 외

Artificial intelligence (AI) is transforming scientific research, including proteomics. Advances in mass spectrometry (MS)-based proteomics data quality, diversity, and scale, combined with groundbreaking AI techniques, …

Diversity