writing
The dangerous answer isn't “I don't know.” It's “$0.”
A column called free, a year of NULLs, and a query that reported a year of revenue as zero without a single error. Why reviewing SQL can't catch a silent error — and what does.
A wrong number doesn’t error. It renders. It gets pasted into a board deck and argued from for weeks, because nothing about it looks broken — and numbers, once they are on a chart, are believed.
This is the story of one wrong number. We tell it because it is the reason our product exists, and because once you have seen the shape of it, you will recognise the same shape in every AI analyst demo you watch.
The column
Evalyst started inside a usage-billing video platform. In its production database there is
a column called free that marks whether a stream was billable. It has been there for
years, and it looks as boring as a column can look.
For every row written before 2023, free is NULL — not 0, not 1 — because the flag
was added later and nobody backfilled the history. Nobody backfills the history. There was
always a more urgent migration, and the people who knew the flag’s birthday carried the
fact around in their heads, where it worked fine.
Facts that live in heads work fine until something that has no access to heads starts writing queries.
The query
Ask a competent text-to-SQL tool for last year’s live-streaming revenue and it writes something like:
-- looks right. drops every pre-2023 row.
SELECT SUM(amount) FROM streams
WHERE free = 0 AND year = 2025;
WHERE free = 0 is the obvious reading of that schema. It is what most humans would write
on their first day. It is also wrong, for a reason every database course mentions and
every schema eventually weaponises: NULL = 0 is never true. Every historical row
silently drops out of the sum, and an entire year of revenue comes back as zero.
The query ran. Nothing errored. The chart rendered.
Why nobody catches it
The instinct is to review the SQL. But the SQL is the one thing here that looks right — a reviewer checks the join, the date range, the aggregate, nods, and approves the exact query above. The defect isn’t in the query; it’s in the distance between what the schema says and what the data means, and that distance is invisible in the query text.
Every answer an analyst — human or machine — produces lands in one of four states:
- correct — the figure matched reality.
- abstained — it declined to answer. Underrated, and the behaviour most tools lack.
- wrong, visibly — it produced a figure you would catch. Costs you an afternoon.
- a silent error — it produced a figure, confidently, with no signal anything was off, and the figure was wrong.
The first three are survivable. The fourth is the one that gets pasted into the deck, because it is the only state that never announces itself. And the uncomfortable part is that making the analyst smarter doesn’t remove the fourth state — a better model writes more plausible queries, which makes its silent errors harder to spot, not rarer in consequence.
What actually catches it
You catch a silent error the way engineering has always caught regressions: with a test that already knows the answer.
Somebody at the company has already signed off on last year’s revenue. That number exists. The fix is to write it down as ground truth — a bank of questions with answers a human confirmed — and to re-run the analyst against that bank every single time the definition of a metric changes. Not once at setup. Every time, because meaning drifts: someone edits what “revenue” includes, a flag gets a third value, a backfill lands.
Then you make the bank load-bearing: if a change to what your metrics mean produces even one silent error against the confirmed answers, the change does not ship. No judgment call, no “it’s probably fine.” The gate either passes or the dashboards keep their old definitions.
That is the whole idea. It is not sophisticated. It is a regression suite for meaning — and it is the difference between an analyst you audit once and an analyst that is audited continuously, by the numbers your own team vouched for.
The question to ask any AI analyst
If you are evaluating one — ours or anyone’s — the demo will show you the happy path: a plain-language question, a confident answer, a chart. Ask instead:
“Which of the four states is this answer in, and how do you know?”
If the honest answer is “we don’t know, but the model is very good,” you are being asked to believe, not to verify. The model may well be very good. So was the query above.
We built Evalyst around this gate: it connects to your databases read-only, drafts what your data means for your team to correct, and turns your confirmed answers into the bank that stands between every change and every dashboard. Each source carries a scorecard — accuracy, silent-error count, bank size, last run — measured from real runs, never asserted. If your revenue lives in a database and its meaning lives in someone’s head, we’d like to talk.