Identifying Consistent Statements about Numerical Data with Dispersion-Corrected Subgroup Discovery
Existing algorithms for subgroup discovery with numerical targets do not optimize the error or target variable dispersion of the groups they find. This often leads to unreliable or inconsistent statements about the data, rendering practical applications, especially in scientific domains, futile. Therefore, we here extend the optimistic estimator framework for optimal subgroup discovery to a new class of objective functions: we show how tight estimators can be computed efficiently for all functions that are determined by subgroup size (non-decreasing dependence), the subgroup median value, and a dispersion measure around the median (non-increasing dependence). In the important special case when dispersion is measured using the average absolute deviation from the median, this novel approach yields a linear time algorithm. Empirical evaluation on a wide range of datasets shows that, when used within branch-and-bound search, this approach is highly efficient and indeed discovers subgroups with much smaller errors.
Code (0)
등록된 구현이 없습니다.
Tasks
Subgroup DiscoverySimilar Papers 제목 키워드 기반
On Context-aware Detection of Cherry-picking in News Reporting
Cherry-picking refers to the deliberate selection of evidence or facts that favor a particular viewpoint while ignoring or distorting evidence that supports an opposing perspective. Manually identifying cherry-picked sta…
What we write about when we write about causality: Features of causal statements across large-scale social discourse
Identifying and communicating relationships between causes and effects is important for understanding our world, but is affected by language structure, cognitive and emotional biases, and the properties of the communicat…
Sentiment AnalysisTopic ModelsComputational Identification of Regulatory Statements in EU Legislation
Identifying regulatory statements in legislation is useful for developing metrics to measure the regulatory density and strictness of legislation. A computational method is valuable for scaling the identification of such…
Dependency ParsingSome Reflections on the Set-based and the Conditional-based Interpretations of Statements in Syllogistic Reasoning
Two interpretations about syllogistic statements are described in this paper. One is the so-called set-based interpretation, which assumes that quantified statements and syllogisms talk about quantity-relationships betwe…
Tell, don't show: Declarative facts influence how LLMs generalize
We examine how large language models (LLMs) generalize from abstract declarative statements in their training data. As an illustration, consider an LLM that is prompted to generate weather reports for London in 2050. One…
Fairness