Contributions to Robust and Efficient Methods for Analysis of High-Dimensional Data
A ubiquitous feature of biological data of our era, such as brain functional magnetic resonance imaging or genetic data, is their extra-large sizes and dimensions. However, analyzing such high–dimensional biological data poses significant challenges, since the feature dimension is often much larger than the sample size. This thesis introduces robust and computationally efficient methods to address several common challenges associated with high–dimensional data.In my first manuscript, I propose a coherent approach to variable screening that can accommodate nonlinear associations. I develop a novel variable screening method that transcends traditional linear assumptions by leveraging mutual information, with an intended application in neuroimaging data. This approach allows for a more accurate identification of important variables by capturing nonlinear as well as linear relationships between the outcome and the covariates. This strategy proves to be transformative in the analysis of neuroimaging data, as demonstrated through a detailed examination of the prepossessed Autism Brain Imaging Data Exchange dataset .Then, building on this foundation, I develop new computing techniques for sparse estimation using nonconvex penalties in my second manuscript. These methods address notable challenges in current statistical computing practices, facilitating computationally efficient and robust analyses of complex datasets. While my study in the second manuscript is mainly motivated by computational challenges in sparse estimation using nonconvex penalties, the proposed method can be applied to a considerably general class of optimization problems.In my third manuscript, I contribute to the development of robust modeling of high–dimensional correlated observations by relaxing some of the underlying assumptions for the analysis of such data. I develop a qGaussian linear mixed-effects model, designed to surpass the constraints of conventional Gaussian linear mixed-effects models by accommodating a broader class of distributions that are more robust toward outliers. For correlated observations, this qGaussian model enhances the robustness and flexibility of statistical analyses, providing a more comprehensive tool for modeling the widely-correlated observations frequently encountered in biological and medical studies.Collectively, these contributions aim at addressing the multifaceted challenges of high–dimensional biological data analysis and paving the way for deeper insights into complex biological systems by seamlessly integrating solutions to nonlinearity, nonconvex nonsmooth optimization, and the need for more robust and adaptable models
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis
We describe a novel attribution method which is grounded in Sensitivity Analysis and uses Sobol indices. Beyond modeling the individual contributions of image regions, Sobol indices provide an efficient way to capture hi…
SensitivityExplaining dimensionality reduction results using Shapley values
Dimensionality reduction (DR) techniques have been consistently supporting high-dimensional data analysis in various applications. Besides the patterns uncovered by these techniques, the interpretation of DR results base…
ClusteringDimensionality ReductionBreaking the curse of dimensionality for linear rules: optimal predictors over the ellipsoid
In this work, we address the following question: What minimal structural assumptions are needed to prevent the degradation of statistical learning bounds with increasing dimensionality? We investigate this question in th…
Understanding Layer-wise Contributions in Deep Neural Networks through Spectral Analysis
Spectral analysis is a powerful tool, decomposing any function into simpler parts. In machine learning, Mercer's theorem generalizes this idea, providing for any kernel and input distribution a natural basis of functions…
A Likelihood Ratio Framework for High Dimensional Semiparametric Regression
We propose a likelihood ratio based inferential framework for high dimensional semiparametric generalized linear models. This framework addresses a variety of challenging problems in high dimensional data analysis, inclu…
regressionSelection biasVocal Bursts Intensity Prediction