rizz.dev
homeaboutblogcontactGitHub
  1. Home
  2. /
  3. Glossary
  4. /
  5. sandbagging
Glossary · term165 of 202

sandbagging

A model deliberately underperforming on a capability evaluation to appear less advanced or dangerous than it actually is.

Related terms
autonomy evaluationA test measuring how independently a model can complete multi-step tasks without human guidance, used to compare how close a model's abilities sit to frontier systems.view term->capability evaluationA structured test measuring what a model can actually do, such as autonomous task completion or dangerous skills, used to decide whether it clears a safety bar before release.view term->red-teamingDeliberately probing a model for harmful or unintended behavior before release, using adversarial prompts crafted to surface failures a normal test would miss.view term->
At a glance
cited by
2 posts
categories
2
first used
apr 2026
Appears in
  • The Open Weights Rules Nobody Can Comply WithEvery open weights rule on the table stops working the moment a download finishes. Sort the three asks by when they act and see which ones still hold.opinion1 min
  • Freelance Client Satisfaction: Get Rehired Every TimeBuild freelance client satisfaction that gets you rehired. Written scope, weekly updates, and over-delivery turn one-off projects into retainer income.guides1 min
previoussafety classifiernextsandbox mode