Glossary · term106 of 202
MBPP
A benchmark of short Python programming problems, used to measure how much a compressed prompt inflates a model's output length compared to an uncompressed baseline.
Related terms
GSM8KA benchmark of grade-school math word problems commonly used to test reasoning accuracy under prompt compression.view termHumanEvalA code-generation benchmark where prompts include a function signature, which lets it survive prompt compression better than benchmarks without that structural anchor.view termoutput expansionThe tendency of a model to generate a longer answer once a compressed prompt strips the structural cues that normally signal how terse to be.view term
At a glance
- cited by
- 1 post
- categories
- 1
- first used
- jul 2026