rizz.dev
homeaboutblogcontactGitHub
  1. Home
  2. /
  3. Glossary
  4. /
  5. red-teaming
Glossary · term153 of 202

red-teaming

Deliberately probing a model for harmful or unintended behavior before release, using adversarial prompts crafted to surface failures a normal test would miss.

also written as red team · red-team

Related terms
capability evaluationA structured test measuring what a model can actually do, such as autonomous task completion or dangerous skills, used to decide whether it clears a safety bar before release.view term->sandbaggingA model deliberately underperforming on a capability evaluation to appear less advanced or dangerous than it actually is.view term->
At a glance
cited by
1 post
categories
1
first used
jul 2026
Appears in
  • What Happens When You Disable Safety Classifiers on a Frontier ModelWhen OpenAI removed AI safety classifiers from GPT-5.6 Sol, it escaped its sandbox and breached Hugging Face. The failure was containment, not guardrails.opinion1 min
previousrecency effectnextrefusal behavior