Glossary · term88 of 202
GSM8K
A benchmark of grade-school math word problems commonly used to test reasoning accuracy under prompt compression.
Related terms
BBHBig-Bench Hard, a set of difficult reasoning benchmarks where compressed prompts lose little because the answer is short and reconstructable from intent.view termLLMLinguaA prompt-compression method that reports roughly 20x token reduction at about a 1.5-point accuracy cost, measured on short-answer reasoning benchmarks rather than agent tasks.view termMBPPA benchmark of short Python programming problems, used to measure how much a compressed prompt inflates a model's output length compared to an uncompressed baseline.view term
At a glance
- cited by
- 2 posts
- categories
- 1
- first used
- jul 2026
Token-Compression Skills Do Not Survive MeasurementToken compression skills report savings the invoice never shows. Cache reads carry 87 percent of an agent bill, and cutting tokens can push the cost up.
6 Critics Grade My Post, Then Fix the Skill That Wrote ItSix critics grade the draft, then patch the SKILL.md that wrote it. The merge code, verdict ladder, and guardrails behind a self-improving Claude Code skill.