paper-with-me

홈 › Papers

Target-Independent Active Learning via Distribution-Splitting

2018-09-28 · Xiaofeng Cao, Ivor W. Tsang, Xiaofeng Xu, Guandong Xu

To reduce the label complexity in Agnostic Active Learning (A^2 algorithm), volume-splitting splits the hypothesis edges to reduce the Vapnik-Chervonenkis (VC) dimension in version space. However, the effectiveness of volume-splitting critically depends on the initial hypothesis and this problem is also known as target-dependent label complexity gap. This paper attempts to minimize this gap by introducing a novel notion of number density which provides a more natural and direct way to describe the hypothesis distribution than volume. By discovering the connections between hypothesis and input distribution, we map the volume of version space into the number density and propose a target-independent distribution-splitting strategy with the following advantages: 1) provide theoretical guarantees on reducing label complexity and error rate as volume-splitting; 2) break the curse of initial hypothesis; 3) provide model guidance for a target-independent AL algorithm in real AL tasks. With these guarantees, for AL application, we then split the input distribution into more near-optimal spheres and develop an application algorithm called Distribution-based A^2 (DA^2). Experiments further verify the effectiveness of the halving and querying abilities of DA^2. Contributions of this paper are as follows.

📄 PDF Abstract BibTeX arXiv:1809.10962

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learning

Similar Papers 제목 키워드 기반

Data thinning for convolution-closed distributions

2023-01-18 · Anna Neufeld, Ameer Dharamshi, Lucy L. Gao, Daniela Witten

We propose data thinning, an approach for splitting an observation into two or more independent parts that sum to the original observation, and that follow the same distribution as the original observation, up to a (know…

Model Selection

Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression

2020-05-29 · Domagoj Ćevid, Loris Michel, Jeffrey Näf, Nicolai Meinshausen 외

Random Forest (Breiman, 2001) is a successful and widely used regression and classification algorithm. Part of its appeal and reason for its versatility is its (implicit) construction of a kernel-type weighting function …

regression

ARIS-RSMA Enhanced ISAC System: Joint Rate Splitting and Beamforming Design

2026-02-06 · Xin Jin, Tiejun Lv, Yashuai Cao, Jie Zeng 외 arxiv

This letter proposes an active reconfigurable intelligent surface (ARIS) assisted rate-splitting multiple access (RSMA) integrated sensing and communication (ISAC) system to overcome the fairness bottleneck in multi-targ…

Causal Information Splitting: Engineering Proxy Features for Robustness to Distribution Shifts

2023-05-10 · Bijan Mazaheri, Atalanti Mastakouri, Dominik Janzing, Michaela Hardt

Statistical prediction models are often trained on data from different probability distributions than their eventual use cases. One approach to proactively prepare for these shifts harnesses the intuition that causal mec…

counterfactualfeature selection

Generalized Data Thinning Using Sufficient Statistics

2023-03-22 · Ameer Dharamshi, Anna Neufeld, Keshav Motwani, Lucy L. Gao 외

Our goal is to develop a general strategy to decompose a random variable $X$ into multiple independent random variables, without sacrificing any information about unknown parameters. A recent paper showed that for some w…