Glossary · term118 of 202
NoLiMa
A long-context benchmark that removes lexical overlap between questions and answers, measuring how far model accuracy falls as input length grows past a short baseline.
also written as NoLiMa benchmark
Related terms
F1 scoreA single accuracy metric combining precision and recall, used in some of the studies cited to show how model performance falls as context length increases.view termcontext rotMeasured accuracy decline as a model's input grows longer, even when the extra tokens are irrelevant, showing up well before the advertised context limit is reached.view termeffective lengthThe context length at which a model's benchmark accuracy actually holds up, often far shorter than its advertised maximum window size.view term
At a glance
- cited by
- 1 post
- categories
- 1
- first used
- jul 2026