paper-with-me

홈 › Papers

Beyond the Fold: Quantifying Split-Level Noise and the Case for Leave-One-Dataset-Out AU Evaluation

2026-04-02 · Saurabh Hinduja, Gurmeet Kaur, Maneesh Bilalpur, Jeffrey Cohn, Shaun Canavan arxiv

Subject-exclusive cross-validation is the standard evaluation protocol for facial Action Unit (AU) detection, yet reported improvements are often small. We show that cross-validation itself introduces measurable stochastic variance. On BP4D+, repeated 3-fold subject-exclusive splits produce an empirical noise floor of $\pm 0.065$ in average F1, with substantially larger variation for low-prevalence AUs. Operating-point metrics such as F1 fluctuate more than threshold-independent measures such as AUC, and model ranking can change under different fold assignments. We further evaluate cross-dataset robustness using a Leave-One-Dataset-Out (LODO) protocol across five AU datasets. LODO removes partition randomness and exposes domain-level instability that is not visible under single-dataset cross-validation. Together, these results suggest that gains often reported in cross-fold validation may fall within protocol variance. Leave-one-dataset-out cross-validation yields more stable and interpretable findings

📄 PDF Abstract BibTeX arXiv:2604.02162

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quantifying Extrinsic Curvature in Neural Manifolds

2022-12-20 · Francisco Acosta, Sophia Sanborn, Khanh Dao Duc, Manu Madhav 외

The neural manifold hypothesis postulates that the activity of a neural population forms a low-dimensional manifold whose structure reflects that of the encoded task variables. In this work, we combine topological deep g…

Dimensionality ReductionTopological Data Analysis

Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects

2024-11-19 · Ke Liang Xiao, Noah Marshall, Atish Agarwala, Elliot Paquette

In recent years, signSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consensus that signSGD acts to precondition op…

Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models

2026-07-12 · Jiacheng Zheng, Chang Guo, Zixuan Wang, Xinyu Liu arxiv

Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that res…

Jailbreaking Embodied LLMs via Action-level Manipulation

2026-03-02 · Xinyu Huang, Qiang Yang, Leming Shen, Zijing Ma 외 arxiv

Embodied Large Language Models (LLMs) enable AI agents to interact with the physical world through natural language instructions and actions. However, beyond the language-level risks inherent to LLMs themselves, embodied…

Geometry-Correct Diffusion Posterior Sampling with Denoiser-Pullback Curvature Guidance and Manifold-Aligned Damping

2026-05-27 · Seunghyeok Shin, Minwoo Kim, Dabin Kim, Hongki Lim arxiv

Diffusion posterior sampling conditions diffusion priors on measurements, but data-consistency updates are typically scaled by hand-tuned guidance weights and can destabilize sampling under stiff, operator-dependent curv…

MRI Reconstruction