paper-with-me

홈 › Papers

JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models

2023-11-01 · Mi Zhang, Xudong Pan, Min Yang

In this paper, we present JADE, a targeted linguistic fuzzing platform which strengthens the linguistic complexity of seed questions to simultaneously and consistently break a wide range of widely-used LLMs categorized in three groups: eight open-sourced Chinese, six commercial Chinese and four commercial English LLMs. JADE generates three safety benchmarks for the three groups of LLMs, which contain unsafe questions that are highly threatening: the questions simultaneously trigger harmful generation of multiple LLMs, with an average unsafe generation ratio of $70\%$ (please see the table below), while are still natural questions, fluent and preserving the core unsafe semantics. We release the benchmark demos generated for commercial English LLMs and open-sourced English LLMs in the following link: https://github.com/whitzard-ai/jade-db. For readers who are interested in evaluating on more questions generated by JADE, please contact us. JADE is based on Noam Chomsky's seminal theory of transformational-generative grammar. Given a seed question with unsafe intention, JADE invokes a sequence of generative and transformational rules to increment the complexity of the syntactic structure of the original question, until the safety guardrail is broken. Our key insight is: Due to the complexity of human language, most of the current best LLMs can hardly recognize the invariant evil from the infinite number of different syntactic structures which form an unbound example space that can never be fully covered. Technically, the generative/transformative rules are constructed by native speakers of the languages, and, once developed, can be used to automatically grow and transform the parse tree of a given question, until the guardrail is broken. For more evaluation results and demo, please check our website: https://whitzard-ai.github.io/jade.html.

📄 PDF Abstract BibTeX arXiv:2311.00286

Code (1)

whitzard-ai/jade-db 공식 구현

Tasks

Natural Questions

Similar Papers 제목 키워드 기반

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring

2025-08-28 · Junjie Chu, Mingjie Li, Ziqing Yang, Ye Leng 외 arxiv

Accurately determining whether a jailbreak attempt has succeeded is a fundamental yet unresolved challenge. Existing evaluation methods rely on misaligned proxy indicators or naive holistic judgments. They frequently mis…

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

2026-02-06 · Lanbo Lin, Jiayao Liu, Tianyuan Yang, Li Cai 외 arxiv

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response st…

A Collaborative Jade Recognition System for Mobile Devices Based on Lightweight and Large Models

2025-02-20 · Zhenyu Wang, Wenjia Li, Pengyu Zhu

With the widespread adoption and development of mobile devices, vision-based recognition applications have become a hot topic in research. Jade, as an important cultural heritage and artistic item, has significant applic…

JADE: Corpus for Japanese Definition Modelling

2022-06-01 · LREC 2022 6 · Han Huang, Tomoyuki Kajiwara, Yuki Arase

This study investigated and released the JADE, a corpus for Japanese definition modelling, which is a technique that automatically generates definitions of a given target word and phrase. It is a crucial technique for pr…

Definition Modelling

SyntaxGym: An Online Platform for Targeted Evaluation of Language Models

2020-07-01 · ACL 2020 6 · Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian 외

Targeted syntactic evaluations have yielded insights into the generalizations learned by neural network language models. However, this line of research requires an uncommon confluence of skills: both the theoretical know…

Experimental DesignLanguage ModelingLanguage Modelling