paper-with-me

Papers

Unsolved Problems in ML Safety

2021-09-28 · Dan Hendrycks, Nicholas Carlini, John Schulman, Jacob Steinhardt

Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the technical problems that the field needs to address. We present four problems ready for research, namely withstanding hazards ("Robustness"), identifying hazards ("Monitoring"), reducing inherent model hazards ("Alignment"), and reducing systemic hazards ("Systemic Safety"). Throughout, we clarify each problem's motivation and provide concrete research directions.

📄 PDF Abstract BibTeX arXiv:2109.13916

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Apriori-based Analysis of Learned Helplessness in Mathematics Tutoring: Behavioral Patterns by Level, Intervention, and Outcome

2026-04-29 · John Paul P. Miranda arxiv

This study applied the Apriori algorithm to analyze behavioral interaction patterns associated with learned helplessness (LH) in mathematics tutoring system logs. Interaction data were examined across three dimensions: L…

HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification

2026-03-16 · Erik Y. Wang, Sumeet Motwani, James V. Roggeveen, Eliot Hodges 외 arxiv

Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely…

Safety of Multimodal Large Language Models on Images and Texts

2024-02-01 · Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang 외

Attracted by the impressive power of Multimodal Large Language Models (MLLMs), the public is increasingly utilizing them to improve the efficiency of daily work. Nonetheless, the vulnerabilities of MLLMs to unsafe instru…

Survey

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning

2025-08-11 · Huanyu Liu, Jia Li, Yihong Dong, Chang Yu 외 arxiv

Reinforcement learning with verifiable reward (RLVR) has become a promising paradigm for post-training large language models (LLMs) to improve their reasoning capability. However, when the rollout accuracy is low on hard…

Reinforcement Learning

A Note on Argumentative Topology: Circularity and Syllogisms as Unsolved Problems

2021-02-07 · Wlodek W. Zadrozny

In the last couple of years there were a few attempts to apply topological data analysis to text, and in particular to natural language inference. A recent work by Tymochko et al. suggests the possibility of capturing `t…

Natural Language InferenceTopological Data AnalysisWord Embeddings