paper-with-me

Papers

Algorithms and Statistical Models for Scientific Discovery in the Petabyte Era

2019-11-05 · Brian Nord, Andrew J. Connolly, Jamie Kinney, Jeremy Kubica, Gautaum Narayan, Joshua E. G. Peek, Chad Schafer, Erik J. Tollerud, Camille Avestruz, G. Jogesh Babu, Simon Birrer, Douglas Burke, João Caldeira, Douglas A. Caldwell, Joleen K. Carlberg, Yen-Chi Chen, Chuanfei Dong, Eric D. Feigelson, V. Zach Golkhou, Vinay Kashyap, T. S. Li, Thomas Loredo, Luisa Lucie-Smith, Kaisey S. Mandel, J. R. Martínez-Galarza, Adam A. Miller, Priyamvada Natarajan, Michelle Ntampaka, Andy Ptak, David Rapetti, Lior Shamir, Aneta Siemiginowska, Brigitta M. Sipőcz, Arfon M. Smith, Nhan Tran, Ricardo Vilalta, Lucianne M. Walkowicz, John ZuHone

The field of astronomy has arrived at a turning point in terms of size and complexity of both datasets and scientific collaboration. Commensurately, algorithms and statistical models have begun to adapt --- e.g., via the onset of artificial intelligence --- which itself presents new challenges and opportunities for growth. This white paper aims to offer guidance and ideas for how we can evolve our technical and collaborative frameworks to promote efficient algorithmic development and take advantage of opportunities for scientific discovery in the petabyte era. We discuss challenges for discovery in large and complex data sets; challenges and requirements for the next stage of development of statistical methodologies and algorithmic tool sets; how we might change our paradigms of collaboration and education; and the ethical implications of scientists' contributions to widely applicable algorithms and computational modeling. We start with six distinct recommendations that are supported by the commentary following them. This white paper is related to a larger corpus of effort that has taken place within and around the Petabytes to Science Workshops (https://petabytestoscience.github.io/).

📄 PDF Abstract BibTeX arXiv:1911.02479

Code (0)

등록된 구현이 없습니다.

Tasks

Astronomyscientific discovery

Similar Papers 제목 키워드 기반

Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA

2025-05-09 · Priyanka Ghosh, Kjiersten Fagnan, Ryan Connor, Ravinder Pannu 외

The volume of biological data being generated by the scientific community is growing exponentially, reflecting technological advances and research activities. The National Institutes of Health's (NIH) Sequence Read Archi…

Benchmarkingscientific discovery

Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

2025-11-21 · Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf, Yu Su arxiv

Scientific archives now contain hundreds of petabytes of data across genomics, ecology, climate, and molecular biology that could reveal undiscovered patterns if systematically analyzed at scale. Large-scale, weakly-supe…

IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular Data

2025-10-10 · Tao Feng, Lizhen Qu, Niket Tandon, Gholamreza Haffari arxiv

Causal discovery is fundamental to scientific research, yet traditional statistical algorithms face significant challenges, including expensive data collection, redundant computation for known relations, and unrealistic …

AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing

2025-10-24 · Samuel Bright-Thonney, Christina Reissel, Gaia Grosso, Nathaniel Woodward 외 arxiv

Novelty detection in large scientific datasets faces two key challenges: the noisy and high-dimensional nature of experimental data, and the necessity of making statistically robust statements about any observed outliers…

Dimensionality ReductionData AugmentationAnomaly Detection

A Field Guide to Scientific XAI: Transparent and Interpretable Deep Learning for Bioinformatics Research

2021-10-13 · Thomas P Quinn, Sunil Gupta, Svetha Venkatesh, Vuong Le

Deep learning has become popular because of its potential to achieve high accuracy in prediction tasks. However, accuracy is not always the only goal of statistical modelling, especially for models developed as part of s…

Deep LearningExplainable Artificial Intelligence (XAI)scientific discovery