whitzard-ai/jade-db
"他山之石、可以攻玉":复旦JADE团队发布的大模型测评与治理系列
What it solves
JADE addresses the challenge of evaluating the safety guardrails of Large Language Models (LLMs). It specifically solves the problem of "low trigger rates," where standard safety tests often fail to bypass a model's safety filters, by creating targeted, high-risk test datasets that can effectively probe for vulnerabilities in model alignment.
How it works
The platform uses linguistic mutation to automatically transform low-trigger-rate "seed questions" into high-risk questions. These mutations create natural-sounding text that is more likely to bypass safety filters. The resulting datasets cover four major categories: core values, illegal activities, infringement of rights, and discrimination/bias, spanning over 30 sub-categories.
Who it’s for
It is designed for AI researchers, developers, and safety auditors who need to rigorously test the safety alignment of LLMs, multimodal models, and AI agents to ensure they do not generate harmful or illegal content.
Highlights
- Linguistic Mutation: Automatically generates high-risk prompts from simple seeds to better test model robustness.
- Comprehensive Coverage: Includes datasets for Chinese and English models across various safety domains (e.g., illegal acts, privacy infringement, and bias).
- Multi-version Benchmarks: Provides different difficulty levels, including Easy, Medium, and High-risk (cross-model transferable) datasets.
- Expanded Ecosystem: Includes specialized tools for phone agents (Jade-BadPhoneAgent), multimodal hallucinations (Jade-HAL), and reasoning chain safety (Jade-LRMGuard).
Related
- Project
- Project
- Project
- Dispatch
- Dispatch