paper-with-me

Computer Security

1개 벤치마크 · 논문 70편 · 이 태스크의 논문 보기 →

Benchmarks

BIG-bench

결과 1개

Most implemented

Defending Against Neural Fake News

2019-05-29 · 구현 4개

Agent Safety Should Be a Runtime Contract

2026-08-11 · 구현 2개

Papers

Agent Safety Should Be a Runtime Contract

2026-08-11 · Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang hf

The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate f…

Computer Security

REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version)

2026-04-30 · Jun Yeon Won, Xin Jin, Shiqing Ma, Zhiqiang Lin arxiv

Large Language Models (LLMs) have achieved remarkable progress in recent years, driving their adoption across a wide range of domains, including computer security. In reverse engineering, LLMs are increasingly applied to…

Computer Security

Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents

2026-03-16 · Simone Aonzo, Merve Sahin, Aurélien Francillon, Daniele Perito arxiv

Artificial intelligence (AI) systems are increasingly adopted as tool-using agents that can plan, observe their environment, and take actions over extended time periods. This evolution challenges current evaluation pract…

Computer Security

CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization

2025-10-09 · Debeshee Das, Luca Beurer-Kellner, Marc Fischer, Maximilian Baader arxiv

The increasing adoption of LLM agents with access to numerous tools and sensitive data significantly widens the attack surface for indirect prompt injections. Due to the context-dependent nature of attacks, however, curr…

Computer Security

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

2025-02-24 · Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley 외

We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of pr…

Computer Security

The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI

2025-01-30 · Peter Hall, Olivia Mundahl, Sunoo Park

Calls for transparency in AI systems are growing in number and urgency from diverse stakeholders ranging from regulators to researchers to users (with a comparative absence of companies developing AI). Notions of transpa…

Computer Security

전체 70편 보기 →