paper-with-me

홈 › Papers

Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation

2026-02-23 · Callum Canavan, Aditya Shrivastava, Allison Qi, Jonathan Michala, Fabien Roger arxiv

To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones (easy-to-hard generalization), or using unsupervised training algorithms to steer models with no external labels at all (unsupervised elicitation). Although techniques from both paradigms have been shown to improve model accuracy on a wide variety of tasks, we argue that the datasets used for these evaluations could cause overoptimistic evaluation results. Unlike many real-world datasets, they often (1) have no features with more salience than truthfulness, (2) have balanced training sets, and (3) contain only data points to which the model can give a well-defined answer. We construct datasets that lack each of these properties to stress-test a range of standard unsupervised elicitation and easy-to-hard generalization techniques. We find that no technique reliably performs well on any of these challenges. We also study ensembling and combining easy-to-hard and unsupervised techniques, and find they only partially mitigate performance degradation due to these challenges. We believe that overcoming these challenges should be a priority for future work on unsupervised elicitation.

📄 PDF Abstract BibTeX arXiv:2602.20400

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

2024-04-15 · Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka 외

This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, deve…

Toward an Evaluation Science for Generative AI Systems

2025-03-07 · Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell 외

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly us…

SoK: Towards Security and Safety of Edge AI

2024-10-07 · Tatjana Wingarz, Anne Lauscher, Janick Edinger, Dominik Kaaser 외

Advanced AI applications have become increasingly available to a broad audience, e.g., as centrally managed large language models (LLMs). Such centralization is both a risk and a performance bottleneck - Edge AI promises…

Hybrid Modelling in Oncology: Sucesses, Challenges and Hopes

2019-01-17

In this review we make the statement that hybrid models in oncology are required as a mean for enhanced data integration. In the context of systems oncology, experimental and clinical data need to be at the heart of the …

Data IntegrationDrug Discovery

I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts

2026-06-12 · Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin arxiv

This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspective of safety assurance. This work has built upon the rapid integration of LLMs ac…