01
The ledger is real, and it is ours
Assay's primary corpus is Kernel's own operation — 253,417 spans, 989,489 sealed turns, 3,471 sessions and 30 principals, from 2026-05-14 to 2026-08-04.
02
A pricing gap caught, and corrected without rewriting history
One model sat correctly unpriced for eight days holding $3,567.33, because the ledger was honest and nothing read the bucket. The fix was a tripwire that reads it on a schedule, plus 86,235 append-only correction rows.
Recovered by correction+$5,575.91
Effective, resolved at read$45,260.38
All 86,235 rows stamped assay@6 · measured 2026-08-03 · every original figure still intact beside its correction
03
The dashboard that was lying, caught by its own discipline
Our own report pages summed the raw ledger and showed $37,129.15 where the corrected truth was $42,705.06 — 13.1% low — because their SQL lived outside the fence that polices cost surfaces. The fence was widened, the pages were regenerated from the shared corrected relation, and the surface was ultimately retired.
04
What did the Mac app cost to build? Answered, and then answered better
The naive answer was unavailable: of the 52,335 spans that built one product, 3 carried its project tag. Assay derives it instead, from path references in the harness's own session documents.
3 of 52,335 spans carried the project tag. A number that is unavailable is reported as unavailable; the derivation that replaced it is recorded with its method, its window and its evidence state, so the two answers can be compared instead of one quietly replacing the other.
05
The breakdown reconciles to the ledger
A product breakdown a reader cannot tie back to the total is a breakdown nobody can trust. Summed over 213 products the weighted column comes to $30,795.52 — exactly the covered population, to 123 microUSD of symmetric rounding.
06
Cost-per-outcome became a number for the first time
Before the attribution contract, the obvious join matched 0 of 911 outcomes — not by accident but by construction, because two producers legitimately name the same work differently.
07
Labels that aggregate
A free-text goal judge produced 567 distinct labels across 1,141 sessions, 478 of them used exactly once — so one real goal split across a dozen rows and a dozen spend totals. A closed, fixture-pinned vocabulary enforced at the parse boundary fixed it.
The bound: this fixed aggregation, not per-session accuracy. On a 7B local judge, 13.7% of answers land on the escape label, and at least 32 of those 135 literally say the context does not say. The remedy is a larger local model — the sovereignty rule constrains where the model runs, not how big it is.
08
The substrate travels
Assay's second installation is a different product, a different producer, and its own isolated store — and it now carries the strongest annotation tier end to end.
09
An upgrade that measured its own assumptions and refused one
A ten-release upgrade landed with zero API breaks. The dry run then showed that a stock repricing pass would have made every spend surface read $10.79 less than the wallet actually charged, because that ledger records revenue at a 2x markup rather than provider cost. The pass was refused and the spec amended in place.