Holdout Set
1개 벤치마크 · 논문 43편 · 이 태스크의 논문 보기 →
Benchmarks
xView3-SAR
Most implemented
Distribution-Free, Risk-Controlling Prediction Sets
Understanding Transformers via N-gram Statistics
Parametric Scaling Law of Tuning Bias in Conformal Prediction
Who's the (Multi-)Fairest of Them All: Rethinking Interpolation-Based Data Augmentation Through the Lens of Multicalibration
Comprehensive dataset of user-submitted articles with ideological and extreme bias from Reddit
TotalVibeSegmentator: Full Body MRI Segmentation for the NAKO and UK Biobank
Papers
ML-Powered LDAP Reconnaissance Detection using Weak Supervision
Lightweight Directory Access Protocol (LDAP) is a protocol that allows users to query and modify Active Directory (AD) data. By default, all users have read access to all AD data through LDAP, making it a common initial …
Holdout SetIn-Domain Supervised Pathology Report Classification: A Reproducible Pipeline from Data Curation to Production-Matched Evaluation
We introduce an in-domain supervised pipeline designed to counter the out-of-distribution performance drop that hampers supervised biomedical NLP models, a problem observed when models trained on pathology reports are mo…
Holdout SetDoing More with Less: Data Augmentation for Sudanese Dialect Automatic Speech Recognition
Although many Automatic Speech Recognition (ASR) systems have been developed for Modern Standard Arabic (MSA) and Dialectal Arabic (DA), few studies have focused on dialect-specific implementations, particularly for low-…
Speech RecognitionData AugmentationHoldout SetAdaptive-Sensorless Monitoring of Shipping Containers
Monitoring the internal temperature and humidity of shipping containers is essential to preventing quality degradation during cargo transportation. Sensorless monitoring -- machine learning models that predict the intern…
Holdout SetGeographic Transferability of Machine Learning Models for Short-Term Airport Fog Forecasting
Short-term forecasting of airport fog (visibility < 1.0 km) presents challenges in geographic generalization because many machine learning models rely on location-specific features and fail to transfer across sites. This…
Feature EngineeringHoldout SetHoldout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
Fine-tuning large pretrained language models is a common approach for aligning them with human preferences, but noisy or off-target examples can dilute supervision. While small, well-chosen datasets often match the perfo…
Holdout Set