paper-with-me

Papers

Zero-Shot Performance Prediction for Probabilistic Scaling Laws

2025-10-19 · Viktoria Schram, Markus Hiller, Daniel Beck, Trevor Cohn arxiv

The prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task's data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to $30$ LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes.

📄 PDF Abstract BibTeX arXiv:2510.16743

Code (0)

등록된 구현이 없습니다.

Tasks

Gaussian ProcessesActive Learning

Similar Papers 제목 키워드 기반

Universal Diffusion-Based Probabilistic Downscaling

2026-02-12 · Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey 외 arxiv

We introduce a universal diffusion-based downscaling framework that lifts deterministic low-resolution weather forecasts into probabilistic high-resolution predictions without any model-specific fine-tuning. A single con…

Weather Forecasting

CLAREL: Classification via retrieval loss for zero-shot learning

2019-05-31 · Boris N. Oreshkin, Negar Rostamzadeh, Pedro O. Pinheiro, Christopher Pal

We address the problem of learning fine-grained cross-modal representations. We propose an instance-based deep metric learning approach in joint visual and textual space. The key novelty of this paper is that it shows th…

ClassificationGeneral ClassificationGeneralized Zero-Shot LearningMetric Learning+3

ZeroPrompt: Scaling Prompt-Based Pretraining to 1,000 Tasks Improves Zero-Shot Generalization

2022-01-18 · Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao 외

We propose a multitask pretraining approach ZeroPrompt for zero-shot generalization, focusing on task scaling and zero-shot prompting. While previous models are trained on only a few dozen tasks, we scale to 1,000 tasks …

Zero-shot GeneralizationZero-Shot Learning

AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models

2026-03-31 · Yubo Cui, Xianchao Guan, Zijun Xiong, Zheng Zhang arxiv

Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fine-tuning methods often disrupt pre-train…

Zero-shot GeneralizationAdversarial Robustness

Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics

2024-06-18 · Hyojin Kim, Jiyoon Lee, Yonghyun Jeong, Haneol Jang 외

This paper presents a novel perspective for enhancing anti-spoofing performance in zero-shot data domain generalization. Unlike traditional image classification tasks, face anti-spoofing datasets display unique generaliz…

Domain GeneralizationFace Anti-Spoofingimage-classificationImage Classification