CosmoFlow: Scale-Aware Representation Learning for Cosmology with Flow Matching
Generative machine learning models have been demonstrated to be able to learn low dimensional representations of data that preserve information required for downstream tasks. In this work, we demonstrate that flow matching based generative models can learn compact, semantically rich latent representations of field level cold dark matter (CDM) simulation data without supervision. Our model, CosmoFlow, learns representations 32x smaller than the raw field data, usable for field level reconstruction, synthetic data generation, and parameter inference. Our model also learns interpretable representations, in which different latent channels correspond to features at different cosmological scales.
Code (0)
등록된 구현이 없습니다.
Tasks
Synthetic Data GenerationRepresentation LearningSimilar Papers 제목 키워드 기반
CosmoFlow: Using Deep Learning to Learn the Universe at Scale
Deep learning is a promising tool to determine the physical model that describes our universe. To handle the considerable computational cost of this problem, we present CosmoFlow: a highly scalable deep learning applicat…
Deep LearningThe Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism
We present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional sample…
2kConsistently Large Cosmic Flows on Scales of 100 Mpc/h: a Challenge for the Standard LCDM Cosmology
Peculiar velocity surveys have non-uniform spatial distributions of tracers, so that the bulk flow estimated from them does not correspond to that of a simple volume such as a sphere. Thus bulk flow estimates are general…
Beyond AI as Assistants: Toward Autonomous Discovery in Cosmology
Recent advances in artificial intelligence (AI) agents are pushing AI beyond tools toward autonomous scientific discovery. We discuss two complementary agentic systems for cosmology: \texttt{CMBEvolve}, which targets tas…
Out-of-Distribution DetectionClairvoyant Prefetching for Distributed Machine Learning I/O
I/O is emerging as a major bottleneck for machine learning training, especially in distributed environments. Indeed, at large scale, I/O takes as much as 85% of training time. Addressing this I/O bottleneck necessitates …
BIG-bench Machine Learning