Generation Is Not Validation

GENERATION IS NOT VALIDATION

Build analytics that's true, not just plausible.

The reliability layer for AI-built analytics.

[01]

The gap

AI writes SQL and dbt faster than anyone can check it. Speed is real. It is not proof.

Plausible code can encode the wrong metric. A passing test can hide a false assumption. A polished dashboard can carry an error all the way to a decision. Fluency is exactly what makes the failure hard to see.

Generating an answer and validating it are different acts, and the gap between them is where wrong numbers reach the people who trust them. Plausible is not true. Confidence must be earned.

[02]

Watch it catch
a wrong number

A correct net-revenue query is allowed. A plausible hallucination, off by $87, is blocked. The right number with no query behind it is blocked too, on governance. Offline, in about a second.

plumbline: make gate
[03]

How it works

01

Govern the definition, not just the query.

02

Evaluate the answer, not the code.

03

Make the context the agent's source of truth.

04

Put a gate between generation and the stakeholder.

+

Verify before you trust, including your own tools.

[04]

By the numbers

One worked example from the kit: net revenue, governed and gated. Gross is what a plausible query returns. Net is the truth. The gap is the whole point.

$0

Net revenue, the truth

$0

Gross, what looks right

$0

The gap the gate caught

~0s

To run the gate, offline

0

Golden cases

1

Gap we found in our own gate

[05]

Who it's for

Mid to senior analytics and data engineers who generate SQL and dbt with AI and want to ship work they would stake their name on.

If a wrong number has ever reached a stakeholder on your watch, or you are quietly afraid it will, this is built for you. Not a faster magic button: the discipline to move from "I can prompt my way to a query" to "I build AI-augmented analytics I'd stake my name on."

[06]

Get the
starter eval
template

A small, free tool that catches wrong answers from AI-generated SQL before they reach a stakeholder. You write the correct answer down once, by hand, and it blocks any query that does not match, with a readable reason and a non-zero exit for CI. Two worked examples, runs offline in minutes, Python and DuckDB only.

The starter eval template plus occasional launch updates. No spam, unsubscribe anytime.

Prefer to just clone it? It is open on GitHub: github.com/plumblinesh/sql-eval-starter

[07]

The reliability kit

The starter teaches one idea: evaluate the answer, not the code. The full Reliability Kit builds on it with the governance axis, a named failure-mode taxonomy, leakage and structural lint, dbt-layer integration, and evaluator calibration.

In progress. Join the list above to hear when it ships.