concepts
The eval gate
A bank of questions your team confirmed, standing between every change to what your metrics mean and every dashboard that reports them.
The gate is the reason this product exists. Everything else — chat, charts, schedules — is available elsewhere. This is the part that decides whether you can believe the output.
Four states, and only one is dangerous
Every answer the agent gives lands in one of four states when scored against a confirmed answer:
- correct — the figure matched, within a tolerance you set.
- abstained — it declined to answer. This is a good outcome when the data can’t support a conclusion, and it is the behaviour most tools lack.
- wrong — it produced a figure and got it visibly wrong, in a way you would catch.
- silent error — it produced a figure, confidently, with no signal that anything was off, and the figure was wrong.
The last one is the only genuinely dangerous state. A wrong answer that announces itself costs you an afternoon. A confident wrong answer gets pasted into a board deck.
What the gate does
Your ground-truth bank is a set of questions with answers a human confirmed. When you change what a metric means, the candidate version is run against that bank — the real path, the same one a user question takes, not a replay of stored SQL — and it goes live only if it comes back with zero silent errors and at least 90% accuracy.
If it fails, the change does not ship, and you get the per-question breakdown of what moved.
Where the bank comes from
The agent proposes questions from your confirmed semantic model and measures each answer, so what you review is a question with a real figure attached. You confirm it, correct the figure, or reject the question. Only confirmed questions count toward the gate — a question the agent invented and graded itself would be a machine marking its own homework.
Good ground-truth questions tend to be the ones you already argue about: a monthly total someone has checked by hand, a per-product split from a board deck, a number with a known gotcha in it.
Two refusals worth knowing about
Reusing a pass. The gate judges the whole bank but only re-executes questions without a reusable pass, because each question is a full agent run and costs real money. A pass is reusable only when it was measured against the same model version and the same question definition — edit the expected answer and the old pass becomes a verdict about a different question.
Re-scoring instead of re-running. If the scorer was wrong — the bar required a literal string
that a correct answer could phrase differently — the honest fix is to replay the stored answer
against the corrected bar, not to spend another agent run. It records wrong as readily as
correct. But an abstention is never re-scored: that verdict keys off the absence of a confident
figure, and stored answers are truncated, so a fabrication past the cut would read as a clean
decline.
What it costs
Each question is a full agent run — it reads the prompt, plans, queries your database several times, and writes an answer. That is the point: anything cheaper would be testing a different system than the one your team uses. Expect roughly a minute per question, and price a gate before you authorise it; the product tells you what a run will cost before it starts.