paper-with-me

홈 › Papers

Bringing a Ruler Into the Black Box: Uncovering Feature Impact from Individual Conditional Expectation Plots

2021-09-06 · Andrew Yeh, Anhthy Ngo

As machine learning systems become more ubiquitous, methods for understanding and interpreting these models become increasingly important. In particular, practitioners are often interested both in what features the model relies on and how the model relies on them--the feature's impact on model predictions. Prior work on feature impact including partial dependence plots (PDPs) and Individual Conditional Expectation (ICE) plots has focused on a visual interpretation of feature impact. We propose a natural extension to ICE plots with ICE feature impact, a model-agnostic, performance-agnostic feature impact metric drawn out from ICE plots that can be interpreted as a close analogy to linear regression coefficients. Additionally, we introduce an in-distribution variant of ICE feature impact to vary the influence of out-of-distribution points as well as heterogeneity and non-linearity measures to characterize feature impact. Lastly, we demonstrate ICE feature impact's utility in several tasks using real-world data.

📄 PDF Abstract BibTeX arXiv:2109.02724

Code (2)

mixerupper/mltools-fi_cate 공식 구현
mixerupper/ice_feature_impact

Methods 이 논문이 사용한 방법론

Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…

Similar Papers 제목 키워드 기반

Accurate Human Gesture Sensing With Coarse-Grained RF Signatures

2019-06-17 · IEEE Access ( Volume: 7 ) 2019 6 · Hongyu Sun, Zheng Lu, Chin-Ling Chen, Jie Cao 외

RF-based gesture sensing and recognition has increasingly attracted intense academic and industrial interest due to its various device-free applications in daily life, such as elder monitoring, mobile games. State-of-the…

RF-based Gesture Recognition

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

2026-01-13 · Yihan Hong, Huaiyuan Yao, Bolin Shen, Wanpeng Xu 외 arxiv

Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the same criteria inconsistently, produce score attributions that are difficu…

Text Generation

RuleR: Improving LLM Controllability by Rule-based Data Recycling

2024-06-22 · Ming Li, Han Chen, Chenguang Wang, Dang Nguyen 외

Despite the remarkable advancement of Large language models (LLMs), they still lack delicate controllability under sophisticated constraints, which is critical to enhancing their response quality and the user experience.…

Data AugmentationInstruction Following

Reading a Ruler in the Wild

2025-07-09 · Yimu Pan, Manas Mehta, Gwen Sincerbeaux, Jeffery A. Goldstein 외

Accurately converting pixel measurements into absolute real-world dimensions remains a fundamental challenge in computer vision and limits progress in key applications such as biomedicine, forensics, nutritional analysis…

Keypoint Detection

One ruler to measure them all: Benchmarking multilingual long-context language models

2025-03-03 · Yekyung Kim, Jenna Russell, Marzena Karpinska, Mohit Iyyer

We present ONERULER, a multilingual benchmark designed to evaluate long-context language models across 26 languages. ONERULER adapts the English-only RULER benchmark (Hsieh et al., 2024) by including seven synthetic task…

8kAllBenchmarking