Glossary · term30 of 202
capability evaluation
A structured test measuring what a model can actually do, such as autonomous task completion or dangerous skills, used to decide whether it clears a safety bar before release.
also written as capability eval · capability evals
Related terms
Foundation Model Transparency IndexA Stanford scoring project that grades AI model developers on how much they publicly disclose about training data, safety evaluations, and deployment practices.view termautonomy evaluationA test measuring how independently a model can complete multi-step tasks without human guidance, used to compare how close a model's abilities sit to frontier systems.view termfrontier modelOne of the most capable AI models available at a given time, used as the reference point against which smaller or older models get compared.view termred-teamingDeliberately probing a model for harmful or unintended behavior before release, using adversarial prompts crafted to surface failures a normal test would miss.view termsandbaggingA model deliberately underperforming on a capability evaluation to appear less advanced or dangerous than it actually is.view term
At a glance
- cited by
- 1 post
- categories
- 1
- first used
- jul 2026