paper-with-me

홈 › Papers

Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation

2023-11-16 · Yao Lu, Jiayi Wang, Raphael Tang, Sebastian Riedel, Pontus Stenetorp

Recent prompt optimisation approaches use the generative nature of language models to produce prompts -- even rivaling the performance of human-curated prompts. In this paper, we demonstrate that randomly sampling tokens from the model vocabulary as ``separators'' can be as effective as language models for prompt-style text classification. Our experiments show that random separators are competitive baselines, having less than a 1% difference compared to previous self-optimisation methods and showing a 12% average relative improvement over strong human baselines across nine text classification tasks and eight language models. We further analyse this phenomenon in detail using three different random generation strategies, establishing that the language space is rich with potentially good separators, with a greater than 40% average chance that a randomly drawn separator performs better than human-curated separators. These observations challenge the common assumption that an effective prompt should be human readable or task relevant and establish a strong baseline for prompt optimisation research.

📄 PDF Abstract BibTeX arXiv:2311.09569

Code (1)

yaolu/random-prompt 공식 구현 pytorch

Tasks

Language Modellingtext-classificationText Classification

Similar Papers 제목 키워드 기반

Sampling from Stochastic Finite Automata with Applications to CTC Decoding

2019-05-21 · Martin Jansche, Alexander Gutkin

Stochastic finite automata arise naturally in many language and speech processing tasks. They include stochastic acceptors, which represent certain probability distributions over random strings. We consider the problem o…

Unbabel's Submission to the WMT2019 APE Shared Task: BERT-based Encoder-Decoder for Automatic Post-Editing

2019-05-30 · WS 2019 8 · António V. Lopes, M. Amin Farajian, Gonçalo M. Correia, Jonay Trenous 외

This paper describes Unbabel's submission to the WMT2019 APE Shared Task for the English-German language pair. Following the recent rise of large, powerful, pre-trained models, we adapt the BERT pretrained model to perfo…

Automatic Post-EditingDecoderMachine TranslationNMT+1

UncertaintyPlayground: A Fast and Simplified Python Library for Uncertainty Estimation

2023-10-23 · Ilia Azizi

This paper introduces UncertaintyPlayground, a Python library built on PyTorch and GPyTorch for uncertainty estimation in supervised learning tasks. The library offers fast training for Gaussian and multi-modal outcome d…

CPUGPUPrediction Intervals

Language Identification in Code-Switched Text Using Conditional Random Fields and Babelnet

2016-11-01 · WS 2016 11 · Utpal Kumar Sikdar, Bj{\"o}rn Gamb{\"a}ck
Language Identification

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

2026-05-18 · Ziwei Wang, Jing Chen, Ruichao Liang, Zhi Wang 외 arxiv

Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretabil…