Glossary · term08 of 202
abliteration
A technique that locates the internal directions in a model's weights responsible for refusal behavior and edits them out, stripping safety guardrails without retraining.
also written as abliterate · abliterated
Related terms
bypass rateThe percentage of attempts that successfully get a model to ignore its safety restrictions, used as a measure of how effective a jailbreak or guardrail-removal technique is.view termcheckpointA saved snapshot of a model's weights at a specific point in training or fine-tuning, the exact file that gets tested, released, or further modified.view termopen-weight modelAn AI model whose trained parameters are published for anyone to download and run on their own hardware, unlike a closed model accessed only through a hosted API.view termrefusal behaviorThe trained tendency of a model to decline certain requests, implemented as learned patterns inside the weights rather than a hardcoded rule, which makes it removable after release.view termsafety classifierA separate filtering system that screens a model's inputs or outputs for harmful content, distinct from the model's own trained behavior and can be disabled independently.view term
At a glance
- cited by
- 1 post
- categories
- 1
- first used
- jul 2026