rizz.dev
homeaboutblogcontactGitHub
  1. Home
  2. /
  3. Glossary
  4. /
  5. NoLiMa
Glossary · term118 of 202

NoLiMa

A long-context benchmark that removes lexical overlap between questions and answers, measuring how far model accuracy falls as input length grows past a short baseline.

also written as NoLiMa benchmark

Related terms
F1 scoreA single accuracy metric combining precision and recall, used in some of the studies cited to show how model performance falls as context length increases.view term->context rotMeasured accuracy decline as a model's input grows longer, even when the extra tokens are irrelevant, showing up well before the advertised context limit is reached.view term->effective lengthThe context length at which a model's benchmark accuracy actually holds up, often far shorter than its advertised maximum window size.view term->
At a glance
cited by
1 post
categories
1
first used
jul 2026
Appears in
  • Claude Is a Math Problem (And the Context Window Is the Equation)The Claude context window is an equation, not a chat. Why failed attempts anchor the output, where measured degradation starts, and when to clear and re-derive.guides1 min
previousno-op turnnextnpm config set os