Skip to contentWolf-Rayet

Decision records

ADR-0034 A step reports only what its own job measured; committed evidence must say so

Accepted2026-09-02Phase 4, reopened

#Context

Every receipt this system produces is committed. packages/budget/receipts/ holds 30 files in the index; packages/tokens/receipts/ holds 23; packages/eslint-plugin/receipts/ holds 2. That is deliberate and it is what makes a receipt reviewable in a diff.

It also means a workflow step that reads one is making a claim about an event — a render, a probe run, a generation — and nothing inside the file says whether that event happened in this job. Thirteen steps across two workflows carry if: always() so that a diagnostic still prints when the thing it diagnoses fails. On a job where the producing step never ran at all — a failed install, a failed Chromium install, a failed drift check, a cancellation — those steps run anyway, read the checkout's own bytes, and print the same green line they print after a real measurement.

Measured, not reasoned. On the pristine checkout at 57a64b4, with receipts/ exactly as git holds it and no engine having run, every one of the ten reporting steps printed its full verdict and exited 0:

MATCH:  every Scope identical to six decimals across platforms
#scope-fleet       load 0.067731  substrate L 0.842869  calm
03-violations failed on: salience — scope, salience — view
  #scope-queue measures 0.096764 in both 07 and 02; 02 passes, 07 fails
08-undecided passes, warning on #scope-queue
p11-declared-chrome passes: 1 region exempt (146099 px), 1 chrome reinstated by declaration
all four themed variants fail on emission and salience

Not one of those renders happened. The last line asserts four themed variants were measured against their ceilings in a job that opened no browser.

The artifact case is not hypothetical either. budget-receipts from run 33534569376 — a green run — was downloaded and compared file by file against the checkout. Four of its thirty files are byte-identical to the committed ones: themed.json, stability.json, mask-integrity.json and ceiling-derivation.json. The fixtures job produces none of them. By bytes rather than by file count that is the larger half: 1,400,665 of the artifact's 1,836,087 bytes — 76.3% — are the repository, not the run. It is named for a run and published if: always().

The one that fails loudly is the same defect. Same-platform delta vs this runner's themed baseline reads committed receipts/themed.json, which was recorded on darwin-arm64. On a dead linux job it compares those numbers to the linux-x64 record and finds 245 of 246 Scopes moved, worst 1.064e-3, and exits 1 reporting "this platform no longer measures what it recorded." It measured nothing. A false red and a false green are the same fault; which one you get depends only on whose receipts happen to be committed.

Why "did the file change?" is not the test. The obvious guard is to check whether the receipt differs from HEAD. It cannot work here, because this system's central claim is that a platform reproduces its own record *exactly* — a successful run leaves every receipt byte-identical, as it did this pass on all 25. "Unmodified" cannot separate "not produced" from "produced and identical". The only thing that can is the producer saying so.

#Decision

A step that reports on work which did not run in the current job must either not run, or must state that its evidence is committed rather than measured. Printing a verdict is reserved for evidence the job produced.

The mechanism is an evidence stamp. tools/evidence.js exposes markMeasured(files), which a producer calls at the moment it writes a receipt, recording the path in .evidence.json at the repository root. That file is untracked by construction — a fresh checkout has none, so every receipt in it reads as committed, which is what it is. The stamp is additive, so a job that ran the engine and the probes carries both and a job that ran only the engine carries only the engine's. A producer that ran and wrote a file stamps it whether or not it then exits non-zero: a failing engine still measured, and its receipts are still this run's evidence, which is the case if: always() exists for and the one that must keep working.

Reporting steps consume it two ways.

  • Gate. node tools/evidence.js --what <label> --needs <file>… -- <command> runs the command only if every named file was stamped this run. Otherwise it prints EVIDENCE: COMMITTED, NOT MEASURED, names the files, says no verdict is given, and exits 0 without spawning anything — the job is already red from whatever killed the producer, and this step does not add a second cause or a false verdict. Tools that own their own reading — compare-baseline.js, compare-themed-baseline.js — call requireMeasured() directly and return before asserting.
  • Note. node tools/evidence.js --note <file> --covers <path>… always exits 0 and writes PROVENANCE.txt marking every file MEASURED or COMMITTED. Uploaded alongside each artifact, so the artifact carries the statement rather than the reader supplying it. The fixtures upload path is also narrowed to the files that job actually writes.

#The enumeration this closes

Read out of the workflows, not from any instance named in advance. The test applied: *can this step run and print a verdict when its producer did not run, and is its evidence a committed file rather than something the run produced?*

#Workflow / jobStepRuns becauseOn a dead job it asserts
1budget / fixturesSame-platform delta vs this runner's baselineif: always()MATCH: every Scope identical to six decimals, exit 0 — committed receipts against the committed baseline, every darwin value printed under the label linux-x64
2budget / fixturesConfirm 04-calm-light passes, calm, on a light substrateif: always()three Scope loads and substrate lightnesses, each calm, exit 0 — a light-substrate render that did not occur
3budget / fixturesConfirm 03-violations still fails, and on the right ruleif: always()failed on: salience — scope, salience — view with the offending elements, exit 0
4budget / fixturesConfirm 07-under-declared fails on the share check and nothing elseif: always()the share failure, its margin, and #scope-queue measures 0.096764 in both 07 and 02 — a claim that two scenes rasterised identically in this run
5budget / fixturesConfirm 08-undecided passes and carries the undecided warningif: always()08-undecided passes, warning on #scope-queue, with both level-3 shares and the spread
6budget / fixturesConfirm content laundering is caught, and its honest twin is notif: always(), while its producer has no if: at allp4-laundering failed on: slot contract, p11 passes: 1 region exempt (146099 px) — eleven probes reported, zero run
7budget / fixturesUpload receiptsif: always(), whole-directory pathan artifact named for the run. Verified on a green run: 4 of 30 files were the checkout's, none produced by that job
8budget / themedSame-platform delta vs this runner's themed baselineif: always()on linux, DIFFER, 245 Scopes listed as moved, worst 1.064e-3, FAIL: this platform no longer measures what it recorded, exit 1 — a false red; on darwin the same step prints MATCH, exit 0
9budget / themedConfirm every themed 03-violations fails on both checksif: always()four variants listed with their checks and all four themed variants fail on emission and salience
10budget / themedConfirm the overlay violation fails on emission and nothing elseif: always()failed on: emission load, Scope #scope-confirm measured 0.31317
11budget / themedConfirm the legal overlay passes in every theme it is authored forif: always()four scrim loads against four ceilings, all passing
12budget / themedUpload themed receiptsif: always()themed.json and ceiling-derivation.json as this run's output
13budget / contractsUpload the pairing matrixif: always()the committed matrix, published — in its own comment — as "a receipt on every run"

Two more sit next to the class and are treated differently, because the difference is real:

#Workflow / jobStepWhy it is not the same fault
14tokens / rampsConfirm field-night is hue-locked and lowest in the systemIts evidence is committed on every run, not only dead ones: node src/cli.js check compares a fresh generation in memory and writes nothing, so receipts/ramps-*.json are always the checkout's. Those files are the shipped artifact this step exists to judge, and it does judge them, in this run. What a green line here does not establish — and cannot, when apca-selftest or check died before it under if: always() — is that the generator still produces them. It now says so.
15tokens / rampsUpload receiptspackages/tokens/receipts/ and css/ are never written by this job, so token-receipts has never contained anything a run measured. Permanently COMMITTED, and the note now says it rather than leaving the artifact's name to imply otherwise.

core.yml, docs.yml and content-lint.yml carry no if: always() step and no step whose evidence is a committed receipt — every check in them runs its own producer in-process. Re-derive the ceilings and assert config.js still agrees reads committed token receipts but performs in-run the derivation it reports, and has no if: always(), so it is not in the class either.

#Proof

The workflow's own step definitions are read out of budget.yml and executed verbatim, in two states.

Dead job — receipts/ restored to the checkout's bytes, .evidence.json absent, no producer run: 10 of 10 reporting steps printed EVIDENCE: COMMITTED, NOT MEASURED, 0 verdict lines, 0 non-zero exits. Each named the files it would have read.

Restored — run, gameability and themed executed, stamp present, the identical harness re-run: 0 disclosures, 9 verdict lines, 0 non-zero exits. MATCH on both comparators, every confirm step reporting as before.

The unbaselined-platform path was proved the same way, by renaming this platform out of baseline.json: engine exit 1, comparator exit 1, --fill-baseline exit 1 with the file untouched, and the file byte-identical after restore.

#Rejected options

Drop if: always() so the steps simply do not run when the producer failed. One line each, no new file, and it is the "must not run" branch of the rule taken literally. It lost because if: always() is load-bearing for the case it was written for: when the engine *runs* and *fails*, it has written fresh receipts, and the diagnostic tables are the whole reason anyone can tell which Scope moved. Removing it would trade a false green for a blind failure, which is a worse trade — the tables exist because "the build is red" and "the build is red because 11-alert-calm #scope-notices moved 1.064e-3" are different amounts of help.

Check whether the receipt differs from HEAD, and disclose if it does not. No new file at all, and it reads like exactly the right question. It is wrong on this repository specifically: an exact same-platform baseline means a *successful* run leaves every receipt byte-identical to the committed one — all 25 did this pass — so the guard would cry wolf on every green local run and stay silent on precisely the dead job it was added for. It fails in the direction that matters.

Have each step re-run its own producer. Guarantees the evidence, needs no stamp, and would make every step self-contained. Rejected on cost and on meaning: the fixtures job would render 25 scenes six more times, and a check that re-measures is no longer checking *the* measurement the job made — it is making a new one, and two of them can disagree.

Uncommit the receipts. The root cause is that a receipt is committed at all; make them artifacts only and no checkout can supply one. It is the cleanest answer to this specific problem and it was rejected because the receipts being reviewable in a diff is a deliberate property of this repository — it is how baseline.json, the drift checks and every "regenerate and assert no drift" step work. Solving a reporting fault by deleting the evidence trail is a bad trade.

Print the disclosure but still run the check, so nothing is lost. Tempting: the reader gets both the caveat and the numbers. Rejected because a verdict printed next to a caveat is still a verdict in the log, in the step's green tick, and in whatever reads the run summary. The failure mode being closed is a reader taking green for measured, and a green step that says "but not really" does not close it.

Exit non-zero on a disclosure. Considered, because it guarantees nobody misses it. Rejected: the job is already red from whatever killed the producer, and adding ten more red steps buries the actual cause under its own consequences. The disclosure is loud enough in the log, and the exit code stays the producer's to set.

#Consequences

A green step now means something narrower and true. Where a verdict prints, the job measured it.

A new producer must stamp what it writes, or every step downstream of it reports nothing. That is one line beside the writeFileSync, and it fails in the safe direction: forgetting it costs a disclosure, never a false verdict.

.evidence.json and PROVENANCE.txt are gitignored and must stay so. A committed stamp would assert exactly what the stamp exists to disprove. This is the single way the mechanism can be defeated, and it is worth naming: the guard's correctness rests on a fresh checkout having no stamp.

The artifacts change shape. budget-receipts no longer carries the four files the fixtures job never produced, and every artifact carries a PROVENANCE.txt. Anything downstream reading these artifacts by directory listing sees one extra file and, in the fixtures case, four fewer.

This does not make CI green or red. No check's pass condition changed. What changed is which checks are allowed to speak.