Proof

The dated record

Assay's primary corpus is our own operation. Every figure below was measured read-only against that store on the date it carries. Marketing copy drifts from live data, so each number is dated exactly as the product dates its own.

01

The ledger is real, and it is ours

Assay's primary corpus is Kernel's own operation — 253,417 spans, 989,489 sealed turns, 3,471 sessions and 30 principals, from 2026-05-14 to 2026-08-04.

$45,260.38

effective spend, correction-aware, as of 2026-08-03. We are the first customer, and the store is ours.

3,471 sessions · 30 principals

every one of them attributed from principal to session to task to tool call.

02

A pricing gap caught, and corrected without rewriting history

One model sat correctly unpriced for eight days holding $3,567.33, because the ledger was honest and nothing read the bucket. The fix was a tripwire that reads it on a schedule, plus 86,235 append-only correction rows.

Raw ledger$39,684.47
Recovered by correction+$5,575.91
Effective, resolved at read$45,260.38

All 86,235 rows stamped assay@6 · measured 2026-08-03 · every original figure still intact beside its correction

03

The dashboard that was lying, caught by its own discipline

Our own report pages summed the raw ledger and showed $37,129.15 where the corrected truth was $42,705.06 — 13.1% low — because their SQL lived outside the fence that polices cost surfaces. The fence was widened, the pages were regenerated from the shared corrected relation, and the surface was ultimately retired.

−13.1%

found and published, 2026-08-01. The product's own docs say this shipped wrong, here is what it cost, here is the fence that now catches it.

$37,129.15 → $42,705.06

what the retired pages reported, beside what the corrected relation actually held.

04

What did the Mac app cost to build? Answered, and then answered better

The naive answer was unavailable: of the 52,335 spans that built one product, 3 carried its project tag. Assay derives it instead, from path references in the harness's own session documents.

$1,361.33

under the first counting unit.

$740.29

under the corrected one — and the earlier derivation stays fully reproducible, verified byte-identical across all 213 products.

3 of 52,335 spans carried the project tag. A number that is unavailable is reported as unavailable; the derivation that replaced it is recorded with its method, its window and its evidence state, so the two answers can be compared instead of one quietly replacing the other.

05

The breakdown reconciles to the ledger

A product breakdown a reader cannot tie back to the total is a breakdown nobody can trust. Summed over 213 products the weighted column comes to $30,795.52 — exactly the covered population, to 123 microUSD of symmetric rounding.

$30,795.52

the weighted column, summed over 213 products.

123 µUSD

the entire residual — symmetric rounding, on a $30k reconciliation.

06

Cost-per-outcome became a number for the first time

Before the attribution contract, the obvious join matched 0 of 911 outcomes — not by accident but by construction, because two producers legitimately name the same work differently.

892 of 932

outcomes carry a recorded binding — 95.7%. Of those, 309 bound to real spend, 583 assert a measured zero, and 40 are honestly unknown.

0

spans claimed by two beats. The measured zeros stay in the denominator; only the genuine unknowns are excluded, and their count is printed beside every ratio.

07

Labels that aggregate

A free-text goal judge produced 567 distinct labels across 1,141 sessions, 478 of them used exactly once — so one real goal split across a dozen rows and a dozen spend totals. A closed, fixture-pinned vocabulary enforced at the parse boundary fixed it.

567 → 31

distinct labels. Churn fell from 49.7% to 2.7%.

478 → 3

singletons — labels used exactly once, which is the shape of an aggregate that cannot aggregate.

The bound: this fixed aggregation, not per-session accuracy. On a 7B local judge, 13.7% of answers land on the escape label, and at least 32 of those 135 literally say the context does not say. The remedy is a larger local model — the sovereignty rule constrains where the model runs, not how big it is.

08

The substrate travels

Assay's second installation is a different product, a different producer, and its own isolated store — and it now carries the strongest annotation tier end to end.

625 spans · 3,033 turns

across 59 principal keys, in a store that is not ours and never touches ours.

4 declared goals

producer-declared session goals written by the live app on 2026-08-03 — the strongest annotation tier, sourced rather than inferred.

09

An upgrade that measured its own assumptions and refused one

A ten-release upgrade landed with zero API breaks. The dry run then showed that a stock repricing pass would have made every spend surface read $10.79 less than the wallet actually charged, because that ledger records revenue at a 2x markup rather than provider cost. The pass was refused and the spec amended in place.

The spec predicted $0.00.

The measurement contradicted it. The measurement won.

$10.79

the drift a stock pass would have introduced — caught in a dry run, before a single row was written.

Fidelity stamp

All figures on this page were measured read-only against the live store on 2026-08-03, except where a different date is named beside the number. The window runs 2026-05-14 to 2026-08-04. Assay 0.20.0, priced against assay@6 and later.

These numbers do not update.

Every number, with its residual attached.

The rules those figures are held to are numbered, versioned, and enforced by tests rather than by aspiration.