paper-with-me

홈 › Papers

PERC: a suite of software tools for the curation of cryoEM data with application to simulation, modelling and machine learning

2025-03-17 · Beatriz Costa-Gomes, Joel Greer, Nikolai Juraschko, James Parkhurst, Jola Mirecka, Marjan Famili, Camila Rangel-Smith, Oliver Strickson, Alan Lowe, Mark Basham, Tom Burnley

Ease of access to data, tools and models expedites scientific research. In structural biology there are now numerous open repositories of experimental and simulated datasets. Being able to easily access and utilise these is crucial for allowing researchers to make optimal use of their research effort. The tools presented here are useful for collating existing public cryoEM datasets and/or creating new synthetic cryoEM datasets to aid the development of novel data processing and interpretation algorithms. In recent years, structural biology has seen the development of a multitude of machine-learning based algorithms for aiding numerous steps in the processing and reconstruction of experimental datasets and the use of these approaches has become widespread. Developing such techniques in structural biology requires access to large datasets which can be cumbersome to curate and unwieldy to make use of. In this paper we present a suite of Python software packages which we collectively refer to as PERC (profet, EMPIARreader and CAKED). These are designed to reduce the burden which data curation places upon structural biology research. The protein structure fetcher (profet) package allows users to conveniently download and cleave sequences or structures from the Protein Data Bank or Alphafold databases. EMPIARreader allows lazy loading of Electron Microscopy Public Image Archive datasets in a machine-learning compatible structure. The Class Aggregator for Key Electron-microscopy Data (CAKED) package is designed to seamlessly facilitate the training of machine learning models on electron microscopy data, including electron-cryo-microscopy-specific data augmentation and labelling. These packages may be utilised independently or as building blocks in workflows. All are available in open source repositories and designed to be easily extensible to facilitate more advanced workflows if required.

📄 PDF Abstract BibTeX arXiv:2503.13329

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Methods 이 논문이 사용한 방법론

AlphaFold 설명 없음

Similar Papers 제목 키워드 기반

Agptools: a utility suite for editing genome assemblies

2025-03-26 · Edward S. Ricemeyer, Rachel A. Carroll, Wesley C. Warren

The AGP format is a tab-separated table format describing how components of a genome assembly fit together. A standard submission format for genome assemblies is a fasta file giving the sequence of contigs along with an …

AutoCure: Automated Tabular Data Curation Technique for ML Pipelines

2023-04-26 · Mohamed Abdelaal, Rashmi Koparde, Harald Schoening

Machine learning algorithms have become increasingly prevalent in multiple domains, such as autonomous driving, healthcare, and finance. In such domains, data preparation remains a significant challenge in developing acc…

Autonomous DrivingData Augmentation

Adaptive Testing for LLM-Based Applications: A Diversity-based Approach

2025-01-23 · Juyeon Yoon, Robert Feldt, Shin Yoo

The recent surge of building software systems powered by Large Language Models (LLMs) has led to the development of various testing frameworks, primarily focused on treating prompt templates as the unit of testing. Despi…

Diversity

Review of Deep Learning Applications to Structural Proteomics Enabled by Cryogenic Electron Microscopy and Tomography

2025-07-25 · Brady K. Zhou, Jason J. Hu, Jane K. J. Lee, Z. Hong Zhou 외 arxiv

The past decade's "cryoEM revolution" has produced exponential growth in high-resolution structural data through advances in cryogenic electron microscopy (cryoEM) and tomography (cryoET). Deep learning integration into …

CaTE Data Curation for Trustworthy AI

2025-08-20 · Mary Versa Clemens-Sewall, Christopher Cervantes, Emma Rafkin, J. Neil Otte 외 arxiv

This report provides practical guidance to teams designing or developing AI-enabled systems for how to promote trustworthiness during the data curation phase of development. In this report, the authors first define data,…