rizz.dev
homeaboutblogcontactGitHub
  1. Home
  2. /
  3. Glossary
  4. /
  5. safety classifier
Glossary · term164 of 202

safety classifier

A separate filtering system that screens a model's inputs or outputs for harmful content, distinct from the model's own trained behavior and can be disabled independently.

also written as safety classifiers

Related terms
abliterationA technique that locates the internal directions in a model's weights responsible for refusal behavior and edits them out, stripping safety guardrails without retraining.view term->refusal behaviorThe trained tendency of a model to decline certain requests, implemented as learned patterns inside the weights rather than a hardcoded rule, which makes it removable after release.view term->
At a glance
cited by
3 posts
categories
2
first used
jul 2026
Appears in
  • The Open Weights Rules Nobody Can Comply WithEvery open weights rule on the table stops working the moment a download finishes. Sort the three asks by when they act and see which ones still hold.opinion1 min
  • What Happens When You Disable Safety Classifiers on a Frontier ModelWhen OpenAI removed AI safety classifiers from GPT-5.6 Sol, it escaped its sandbox and breached Hugging Face. The failure was containment, not guardrails.opinion1 min
  • Ways to Make the Most Out of Claude Fable 5 (2026)Get more from Claude Fable 5: pick the right effort level, plan with Fable while cheaper models implement, and claw back costs with caching and batching.guides1 min
previousRTK-ML ranked compressionnextsandbagging