paper-with-me

홈 › Papers

TriSig: Assessing the statistical significance of triclusters

2023-06-01 · Leonardo Alexandre, Rafael S. Costa, Rui Henriques

Tensor data analysis allows researchers to uncover novel patterns and relationships that cannot be obtained from matrix data alone. The information inferred from the patterns provides valuable insights into disease progression, bioproduction processes, weather fluctuations, and group dynamics. However, spurious and redundant patterns hamper this process. This work aims at proposing a statistical frame to assess the probability of patterns in tensor data to deviate from null expectations, extending well-established principles for assessing the statistical significance of patterns in matrix data. A comprehensive discussion on binomial testing for false positive discoveries is entailed at the light of: variable dependencies, temporal dependencies and misalignments, and \textit{p}-value corrections under the Benjamini-Hochberg procedure. Results gathered from the application of state-of-the-art triclustering algorithms over distinct real-world case studies in biochemical and biotechnological domains confer validity to the proposed statistical frame while revealing vulnerabilities of some triclustering searches. The proposed assessment can be incorporated into existing triclustering algorithms to mitigate false positive/spurious discoveries and further prune the search space, reducing their computational complexity. Availability: The code is freely available at https://github.com/JupitersMight/TriSig under the MIT license.

📄 PDF Abstract BibTeX arXiv:2306.00643

Code (1)

jupitersmight/trisig 공식 구현

Similar Papers 제목 키워드 기반

Triclustering of Gene Expression Microarray data using Evolutionary Approach

2018-05-14 · Shreya Mishra, Swati Vipsita

In Tri-clustering, a sub-matrix is being created, which exhibit highly similar behavior with respect to genes, conditions and time-points. In this technique, genes with same expression values are discovered across some f…

Clustering

Using Score Distributions to Compare Statistical Significance Tests for Information Retrieval Evaluation

2019-01-30 · Parapar Javier, Losada David E., Presedo-Quindimil Manuel A., Barreiro Alvaro

Statistical significance tests can provide evidence that the observed difference in performance between two methods is not due to chance. In Information Retrieval, some studies have examined the validity and suitability …

Information RetrievalRetrieval

Post-Transfer Learning Statistical Inference in High-Dimensional Regression

2025-04-25 · Nguyen Vu Khai Tam, Cao Huyen My, Vo Nguyen Le Duy

Transfer learning (TL) for high-dimensional regression (HDR) is an important problem in machine learning, particularly when dealing with limited sample size in the target task. However, there currently lacks a method to …

feature selectionregressionTransfer Learningvalid

Computing Valid p-value for Optimal Changepoint by Selective Inference using Dynamic Programming

2020-02-21 · NeurIPS 2020 12 · Vo Nguyen Le Duy, Hiroki Toda, Ryota Sugiyama, Ichiro Takeuchi

There is a vast body of literature related to methods for detecting changepoints (CP). However, less attention has been paid to assessing the statistical reliability of the detected CPs. In this paper, we introduce a nov…

Computational Efficiencyvalid

Towards Statistically Significant Taxonomy Aware Co-location Pattern Detection

2024-06-29 · Subhankar Ghosh, Arun Sharma, Jayant Gupta, Shashi Shekhar

Given a collection of Boolean spatial feature types, their instances, a neighborhood relation (e.g., proximity), and a hierarchical taxonomy of the feature types, the goal is to find the subsets of feature types or their…