paper-with-me

Papers

Identifying OCRs in cfDNA WGS Data by Correlation Clustering

2022-02-19 · Farshad Noravesh, Fahimeh Palizban

In the recent decade, the emergence of liquid biopsy has significantly improved cancer monitoring and detection. Dying cells, including those originating from tumors, shed their DNA into the bloodstream and contribute to a pool of circulating fragments called cell-free DNA (cfDNA). Identifying the tissue origin of these DNA fragments from their epigenetic features has implications in various clinical contexts. Open chromatin regions (OCRs) are important epigenetic features of DNA that reflect cell types of origin. Profiling these features by DNase-seq, ATAC-seq, and histone ChIP-seq provides insights into tissue-specific and disease-specific regulatory mechanisms. Integration of genomic and epigenomic features for cancer detection by liquid biopsy has previously been reported. However, many multimodal analyses require large amounts of cfDNA input and/or multiple types of experiments to cover the genomic and epigenomic aspects of a single sample which is cost and time prohibitive. Thus, methods that capture genomic and epigenomic profiles in a single experiment type with low input requirements are of importance. Predicting OCRs from whole genome sequencing (WGS) data is one such approach. Here, we applied a correlation clustering algorithm to predict OCRs. We used local sequencing depth as input to our algorithm. Multiple processing steps were then applied as follows: count normalization, discrete Fourier transform conversion, graph construction, graph cut optimization by linear programming, and clustering. To validate the proposed method, we compared the output of our predictions (OCR vs. non-OCR) with previously validated open chromatin regions related to human blood samples of the ATAC-db. The percentage of overlap between them is greater than 67%.

📄 PDF Abstract BibTeX arXiv:2202.09618

Code (0)

등록된 구현이 없습니다.

Tasks

Clusteringgraph constructionOptical Character Recognition (OCR)Specificity

Methods 이 논문이 사용한 방법론

BAM Park et al. proposed the bottleneck attention module (BAM), aiming to efficiently improve the representational capability of networks. It uses dilated convolution to enlarge…

Similar Papers 제목 키워드 기반

A DNA Methylation Classification Model Predicts Organ and Disease Site

2025-05-30 · Keng-Jung Lee, Dharanya Sampath, Konstantinos Mavrommatis

Cell-free DNA (cfDNA) analysis is a powerful, minimally invasive tool for monitoring disease progression, treatment response, and early detection. A major challenge, however, is accurately determining the tissue of origi…

Dimensionality ReductionImputation

Computational Methods and Challenges in Cell-Free DNA Analysis for Multi-Cancer Early Detection

2026-06-18 · Nicko Starkey, Marcin W. Wojewodzic, Krzysztof Rzecki arxiv

Cell-free DNA (cfDNA) is a promising avenue for non-invasive multicancer early detection (MCED), in that, it can enable multiple cancer detection simultaneously from a single blood draw, with particular sensitivity to ca…

A Cost Efficient Approach to Correct OCR Errors in Large Document Collections

2019-05-28 · Deepayan Das, Jerin Philip, Minesh Mathew, C. V. Jawahar

Word error rate of an ocr is often higher than its character error rate. This is especially true when ocrs are designed by recognizing characters. High word accuracies are critical to tasks like the creation of content i…

ClusteringLanguage ModellingOptical Character Recognition (OCR)text-to-speech+1

Fully Dynamic Online Selection through Online Contention Resolution Schemes

2023-01-08 · Vashist Avadhanula, Andrea Celli, Riccardo Colini-Baldeschi, Stefano Leonardi 외

We study fully dynamic online selection problems in an adversarial/stochastic setting that includes Bayesian online selection, prophet inequalities, posted price mechanisms, and stochastic probing problems subject to com…

CSTS: A Benchmark for the Discovery of Correlation Structures in Time Series Clustering

2025-05-20 · Isabella Degen, Zahraa S Abdallah, Henry W J Reeve, Kate Robson Brown

Time series clustering promises to uncover hidden structural patterns in data with applications across healthcare, finance, industrial systems, and other critical domains. However, without validated ground truth informat…

ClusteringClustering Algorithms EvaluationClustering Multivariate Time SeriesTime Series+1