paper-with-me

홈 › Papers

Modelling-based experiment retrieval: A case study with gene expression clustering

2015-05-19 · Paul Blomstedt, Ritabrata Dutta, Sohan Seth, Alvis Brazma, Samuel Kaski

Motivation: Public and private repositories of experimental data are growing to sizes that require dedicated methods for finding relevant data. To improve on the state of the art of keyword searches from annotations, methods for content-based retrieval have been proposed. In the context of gene expression experiments, most methods retrieve gene expression profiles, requiring each experiment to be expressed as a single profile, typically of case vs. control. A more general, recently suggested alternative is to retrieve experiments whose models are good for modelling the query dataset. However, for very noisy and high-dimensional query data, this retrieval criterion turns out to be very noisy as well. Results: We propose doing retrieval using a denoised model of the query dataset, instead of the original noisy dataset itself. To this end, we introduce a general probabilistic framework, where each experiment is modelled separately and the retrieval is done by finding related models. For retrieval of gene expression experiments, we use a probabilistic model called product partition model, which induces a clustering of genes that show similar expression patterns across a number of samples. The suggested metric for retrieval using clusterings is the normalized information distance. Empirical results finally suggest that inference for the full probabilistic model can be approximated with good performance using computationally faster heuristic clustering approaches (e.g. $k$-means). The method is highly scalable and straightforward to apply to construct a general-purpose gene expression experiment retrieval method. Availability: The method can be implemented using standard clustering algorithms and normalized information distance, available in many statistical software packages.

📄 PDF Abstract BibTeX arXiv:1505.05007

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringRetrieval

Similar Papers 제목 키워드 기반

Modelling dependency completion in sentence comprehension as a Bayesian hierarchical mixture process: A case study involving Chinese relative clauses

2017-02-02 · Shravan Vasishth, Nicolas Chopin, Robin Ryder, Bruno Nicenboim

We present a case-study demonstrating the usefulness of Bayesian hierarchical mixture modelling for investigating cognitive processes. In sentence comprehension, it is widely assumed that the distance between linguistic …

RetrievalSentence

Evaluation of Information Retrieval Systems Using Structural Equation Modelling

2018-06-25 · Melucci Massimo

The interpretation of the experimental data collected by testing systems across input datasets and model parameters is of strategic importance for system design and implementation. In particular, finding relationships be…

Information RetrievalRetrieval

Modelling Word Burstiness in Natural Language: A Generalised Polya Process for Document Language Models in Information Retrieval

2017-08-20 · Cummins Ronan

We introduce a generalised multivariate Polya process for document language modelling. The framework outlined here generalises a number of statistical language models used in information retrieval for modelling document …

Information RetrievalLanguage ModelingLanguage ModellingRetrieval

Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT

2025-11-17 · Jonathon Dilworth, Hui Yang, Jiaoyan Chen, Yongsheng Gao 외 arxiv

SNOMED CT is a biomedical ontology with a hierarchical representation, modelling terminological concepts at a large scale. Knowledge retrieval in SNOMED CT is critical for its application but often proves challenging due…

Alignment Analysis of Sequential Segmentation of Lexicons to Improve Automatic Cognate Detection

2018-11-20 · ACL 2018 7 · Pranav A

Ranking functions in information retrieval are often used in search engines to recommend the relevant answers to the query. This paper makes use of this notion of information retrieval and applies onto the problem domain…

General ClassificationInformation RetrievalLanguage ModellingRetrieval+1