paper-with-me

Papers

Data Augmentation for Compositional Data: Advancing Predictive Models of the Microbiome

2022-05-20 · Elliott Gordon-Rodriguez, Thomas P. Quinn, John P. Cunningham

Data augmentation plays a key role in modern machine learning pipelines. While numerous augmentation strategies have been studied in the context of computer vision and natural language processing, less is known for other data modalities. Our work extends the success of data augmentation to compositional data, i.e., simplex-valued data, which is of particular interest in the context of the human microbiome. Drawing on key principles from compositional data analysis, such as the Aitchison geometry of the simplex and subcompositions, we define novel augmentation strategies for this data modality. Incorporating our data augmentations into standard supervised learning pipelines results in consistent performance gains across a wide range of standard benchmark datasets. In particular, we set a new state-of-the-art for key disease prediction tasks including colorectal cancer, type 2 diabetes, and Crohn's disease. In addition, our data augmentations enable us to define a novel contrastive learning model, which improves on previous representation learning approaches for microbiome compositional data. Our code is available at https://github.com/cunningham-lab/AugCoDa.

📄 PDF Abstract BibTeX arXiv:2205.09906

Code (1)

cunningham-lab/augcoda 공식 구현

Tasks

Contrastive LearningData AugmentationDisease PredictionRepresentation Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Disentangling Genotype and Environment Specific Latent Features for Improved Trait Prediction using a Compositional Autoencoder

2024-10-25 · Anirudha Powadi, Talukder Zaki Jubery, Michael C. Tross, James C. Schnable 외

This study introduces a compositional autoencoder (CAE) framework designed to disentangle the complex interplay between genotypic and environmental factors in high-dimensional phenotype data to improve trait prediction i…

Diversityregression

TreeMix: Compositional Constituency-based Data Augmentation for Natural Language Understanding

2022-05-12 · NAACL 2022 7 · Le Zhang, Zichao Yang, Diyi Yang

Data augmentation is an effective approach to tackle over-fitting. Many previous works have proposed different data augmentations strategies for NLP, such as noise injection, word replacement, back-translation etc. Thoug…

Constituency ParsingData AugmentationDiversityNatural Language Understanding+2

SUBS: Subtree Substitution for Compositional Semantic Parsing

2022-05-03 · NAACL 2022 7 · Jingfeng Yang, Le Zhang, Diyi Yang

Although sequence-to-sequence models often achieve good performance in semantic parsing for i.i.d. data, their performance is still inferior in compositional generalization. Several data augmentation methods have been pr…

Data AugmentationSemantic Parsing

SUBS: Subtree Substitution for Compositional Semantic Parsing

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Although sequence-to-sequence models often achieve good performance in semantic parsing for i.i.d. data, their performance is still inferior in compositional generalization. Several data augmentation methods have been pr…

Data AugmentationSemantic Parsing

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics

2026-06-10 · Sam Gijsen, Michał Łukomski, Marc-André Schulz, Kerstin Ritter arxiv

Flow matching and diffusion models enable conditional generation across domains ranging from images to proteins, with recent extensions to out-of-distribution contexts. Yet generative models of neural time series have la…

Zero-shot Generalization