Grimoire docs
Live demo: this is your own private copy of a team wiki, and you’re its admin. Edit docs, restore history, accept suggestions, change permissions or import content. Nobody else sees your changes; the copy resets after 3 hours idle.
Docsevaluationevaluation/index.md

Offline harnesses, LLM-as-judge, and regression gates for LLM systems.

publishedevaluationoverviewUpdated 2026-10-05

Evaluation

You cannot improve what you cannot measure, and LLM systems are unusually easy to fool yourself about — a change that looks great on three hand-picked examples can quietly regress on the twenty you didn't check. Evaluation is the discipline of measuring quality systematically so that "it feels better" becomes "recall went from 0.71 to 0.83." It is the difference between engineering an LLM system and tinkering with one.

Documents in this section

Evaluation ties the whole library together: it's how you choose a chunk size, a prompt, or an agent architecture on evidence instead of vibes.