rizz.dev
homeaboutblogcontactGitHub
  1. Home
  2. /
  3. Glossary
  4. /
  5. HumanEval
Glossary · term93 of 202

HumanEval

A code-generation benchmark where prompts include a function signature, which lets it survive prompt compression better than benchmarks without that structural anchor.

Related terms
MBPPA benchmark of short Python programming problems, used to measure how much a compressed prompt inflates a model's output length compared to an uncompressed baseline.view term->output expansionThe tendency of a model to generate a longer answer once a compressed prompt strips the structural cues that normally signal how terse to be.view term->
At a glance
cited by
2 posts
categories
2
first used
apr 2026
Appears in
  • Token-Compression Skills Do Not Survive MeasurementToken compression skills report savings the invoice never shows. Cache reads carry 87 percent of an agent bill, and cutting tokens can push the cost up.meta-analysis1 min
  • Self-Hosting AI for Code: The $0/Month Developer StackBuild a complete self-hosted AI coding stack with Ollama, Continue.dev, and Open WebUI for $0/month. Includes hardware requirements, cost math, and the setup commands.guides1 min
previousHN Algolia APInextinference cost collapse