paper-with-me

홈 › Papers

Scalable Signature-Based Distribution Regression via Reference Sets

2024-10-11 · Andrew Alden, Carmine Ventre, Blanka Horvath

Distribution Regression (DR) on stochastic processes describes the learning task of regression on collections of time series. Path signatures, a technique prevalent in stochastic analysis, have been used to solve the DR problem. Recent works have demonstrated the ability of such solutions to leverage the information encoded in paths via signature-based features. However, current state of the art DR solutions are memory intensive and incur a high computation cost. This leads to a trade-off between path length and the number of paths considered. This computational bottleneck limits the application to small sample sizes which consequently introduces estimation uncertainty. In this paper, we present a methodology for addressing the above issues; resolving estimation uncertainties whilst also proposing a pipeline that enables us to use DR for a wide variety of learning tasks. Integral to our approach is our novel distance approximator. This allows us to seamlessly apply our methodology across different application domains, sampling rates, and stochastic process dimensions. We show that our model performs well in applications related to estimation theory, quantitative finance, and physical sciences. We demonstrate that our model generalises well, not only to unseen data within a given distribution, but also under unseen regimes (unseen classes of stochastic models).

📄 PDF Abstract BibTeX arXiv:2410.09196

Code (0)

등록된 구현이 없습니다.

Tasks

regression

Similar Papers 제목 키워드 기반

Scalable Machine Learning Algorithms using Path Signatures

2025-06-21 · Csaba Tóth

The interface between stochastic analysis and machine learning is a rapidly evolving field, with path signatures - iterated integrals that provide faithful, hierarchical representations of paths - offering a principled a…

Computational EfficiencyGaussian ProcessesTime SeriesTime Series Forecasting

A Fast Text Similarity Measure for Large Document Collections using Multi-reference Cosine and Genetic Algorithm

2018-10-07 · Hamid Mohammadi, Seyed Hossein Khasteh

One of the important factors that make a search engine fast and accurate is a concise and duplicate free index. In order to remove duplicate and near-duplicate documents from the index, a search engine needs a swift and …

Articlestext similarity

A Graphical Model Approach for Matching Partial Signatures

2015-06-01 · CVPR 2015 6 · Xianzhi Du, David Doermann, Wael Abd-Almageed

In this paper, we present a novel partial signature matching method using graphical models. Shape context features are extracted from the contour of signatures to capture local variations, and K-means clustering is used …

Clustering

Modeling Massive Spatial Datasets Using a Conjugate Bayesian Linear Regression Framework

2021-09-09 · Sudipto Banerjee

Geographic Information Systems (GIS) and related technologies have generated substantial interest among statisticians with regard to scalable methodologies for analyzing large spatial datasets. A variety of scalable spat…

Bayesian Inferenceregression

Distribution Regression for Sequential Data

2020-06-10 · Maud Lemercier, Cristopher Salvi, Theodoros Damoulas, Edwin V. Bonilla 외

Distribution regression refers to the supervised learning problem where labels are only available for groups of inputs instead of individual inputs. In this paper, we develop a rigorous mathematical framework for distrib…

regressionTime SeriesTime Series Analysis