Skip to contentWolf-Rayet

Decision records

ADR-0032 A baseline is recorded per platform and gated exactly; the cross-platform bound is retired, not widened

Accepted2026-09-01Phase 4, reopened

#Context

CI run 33528483192 failed the themed job on two Scope values out of 246:

worst per theme:
  field-day        1.064e-3   11-alert-calm #scope-notices
  field-night      4.900e-5   08-undecided #scope-queue
  interior-dark    8.520e-4   17-input-calm #scope-validation
  interior-light   1.040e-3   11-alert-calm #scope-notices

ADR-0004's cross-platform bound is 1.0e-3. Two values were over it, both the same Scope of the same scene, in the two themes on a light substrate.

What that Scope holds that no other does. 11-alert-calm #scope-notices is the only Scope in the themed set carrying three full alert bodies. Across all 83 renders and 246 Scope values, the census is 213 Scopes with no alert, 27 with one, 3 with two, and 3 with three — the three being this Scope in its three themed variants. At 383 characters of rendered prose it is the largest text mass of any Scope in the set, 26% above the next, and alert prose is body copy at level-2 and level-3 weight rather than the short labels the other Scopes carry: 2.7e-6 of drift per character against 1.0e-6 to 1.3e-6 elsewhere in the same scene.

The excess is spread, not one element. Decomposed against the linux-x64 receipts from the failing run:

darwin-arm64linux-x64Δshare of the Scope's drift
#notice-seal0.0215180.021183−3.350e-431.5%
#notice-uplink0.0209760.020572−4.040e-438.0%
#notice-calibration0.0188620.018550−3.120e-429.3%
scope ground0.0091080.009096−1.200e-51.1%

interior-light decomposes to 31.6 / 38.0 / 29.3 / 1.3 — the same proportions. No element is over the bound on its own; the largest is 38% of a number that exceeds it by 6%. The Scope fails because it holds three of them. Nothing is wrong with any one alert, and there is no element to fix.

What the bound was actually measuring. The same Chromium build renders these scenes on both platforms — 151.0.7922.34, from the same pinned Playwright 1.62.1. The only difference is the font rasteriser: CoreText on darwin-arm64, FreeType on linux-x64. Measured across the full set, 245 of 246 Scope values differ between the two platforms. One matches, and it matches by coincidence. The bound was never asserting that the engine returns the same number on every platform; it was asserting that a difference present in effectively every value happened to stay under 1.0e-3. interior-dark was already at 8.520e-4, 85% of it, on a scene nobody had flagged.

What is exact. Same-platform reproducibility is not approximately exact, it is exact, and this was measured on both platforms this pass rather than assumed:

  • darwin-arm64, this machine against its own committed record: 246 of 246 identical, worst delta 0.
  • linux-x64, run 33528483192 against run 33520512425 — two independent runner VMs, 77 minutes apart, on two commits whose difference touches no fixture and no engine file: 246 of 246 identical, worst delta exactly 0.

#Decision

themed-baseline.json holds one record per platform, keyed by ${process.platform}-${process.arch}, each carrying its own node, playwright, chromium and recordedBy alongside its scenes. A run compares against the record for the platform doing the measuring and against nothing else, at REPRODUCIBILITY_TOLERANCE — exact equality. PLATFORM_TOLERANCE no longer governs the themed path.

Two records are committed with this decision. darwin-arm64 is carried through byte-identical from the single record that preceded it — no value re-measured, no evidence spent. linux-x64 is the themed-receipts artifact of run 33528483192, recorded by CI on the platform it describes, at commit c7dca10, whose difference from the tree it is committed against is confined to docs/ and therefore measures these fixture bytes. Its provenance is written into the file rather than left to this record.

A platform with no entry fails and names itself. It does not measure itself and pass. This is ADR-0029's rule for an unbaselined scene, applied to an unbaselined platform, and for the same reason: a first record compared to nothing is a tautology, and a tautology that reports green is worse than a gap that reports red. --write-baseline rewrites only the running platform's entry and prints which other platforms it carried through untouched; --fill-baseline is additive within the running platform's entry and refuses to run at all if that platform has no entry yet.

The gate is exact, and the consequence is accepted deliberately: a runner image that changes its font stack under an unchanged platform name will turn the themed job red. That is the correct outcome. A rasteriser change moves every number in the set, it is exactly the class of change a baseline exists to make visible, and the remedy is a deliberate re-record reviewed in the diff — not a tolerance sized to absorb it in advance.

#Rejected options

Widen PLATFORM_TOLERANCE to admit 1.064e-3. One constant, one line, and the job goes green. It had the merit of being the smallest possible change and of being arguably within the spirit of a bound that was always an estimate. It is rejected on the record, for the reason ADR-0011 and ADR-0030 give about ceilings: a bound raised to admit the value that broke it is not a bound, it is a transcript of the worst thing measured so far. It is also rejected on its own arithmetic. interior-dark sits at 8.520e-4 with no widening at all, so a bound set to clear 1.064e-3 buys margin over one value while another sits at 80% of the new number, and the next prose-heavy component Scope reopens this decision. The quantity being bounded — a difference between two font rasterisers — has nothing to do with the design system and no reason to stay under any number the design system picks.

Per-platform baselines. Chosen. Its cost is stated below.

Change what the scene renders. Dropping one of the three alerts, or shortening their prose, puts #scope-notices under the bound immediately, and it is the only option that addresses the measurement rather than the gate. It lost to the scene's purpose, which its own generator states: fixture 11 is playbook §10's central claim as a measurement — "A monitoring screen where everything is quietly running passes every check" — and the claim is about a screen carrying *three* alerts, none of them demanding attention. Two alerts do not make it. Every alert also carries all three parts of §6's Error pattern because the schema is required rather than encouraged, so the prose is not padding that could be trimmed. Editing the scene to fit the gate is fitting the evidence to the instrument, and it would have left the real finding — that the instrument compares across rasterisers — undiscovered.

Restrict 11-alert-calm out of field-day and interior-light, the way 06-overlay-violation is restricted to one theme. It would have gone green and it uses a mechanism already in the tree. It lost for the reason ADR-0029 gave when rejecting the same move for 14-queue-calm and 27-checkbox-calm: theme restriction exists for scenes built to fail in a theme, and using it to hide a scene that fails for an unrelated reason teaches the reader that a scene's coverage is negotiable. It would also have removed the two renders that produced the entire finding.

Keep one record and gate the ceilings only, not the per-Scope values. The two failing values do not move any derived ceiling, so this would have gone green while leaving the per-Scope table as reporting. It lost to the comment already in cli.js: a Scope can move without moving the worst-Scope number a ceiling is taken from, and that movement is exactly what a baseline exists to catch. It trades the check for the thing the check was for.

Normalise the rasteriser instead — bundle fonts, force a common rasterisation path, or run the measurement in a container on both platforms. This is the option that would preserve a single portable number, and it is the only rejected option that attacks the cause. It lost on cost and on honesty about what is being claimed. The engine measures pixels a browser drew; making darwin and linux draw identical pixels means shipping and pinning a font stack and a rasteriser, and a design system's budget engine that only agrees with itself inside one container has replaced a portable claim with a narrower one while looking like it kept the broad one. It is left open as the thing to build if a portable per-Scope number is ever needed for its own sake, and it is named here so the option is on the record rather than forgotten.

#Consequences

What survives. The claim the suite makes is now: *this platform measures what this platform recorded, exactly.* It is stronger than what it replaced — 0 rather than 1.0e-3 — and it is the claim that catches regressions, because a change to a scene, a token, a component or the engine moves a number on the machine measuring it. It was verified on both platforms this pass, and on linux across two independent runner VMs.

What is lost, plainly. The suite no longer asserts anything about one platform's numbers against another's. There is no committed number that is *the* emission load of a Scope; there are two, and 245 of 246 differ. Anyone reading a value out of themed-baseline.json must now read which platform's record it came from, and any future claim of the form "this Scope's load is X" is a claim about a platform. ADR-0004's Measured-properties table keeps its cross-platform figures as the historical measurement they were, and its bound keeps governing the unthemed run() path against baseline.json, which has not come due. The portable-number claim is not weakened here so much as revealed: it had been false in 245 of 246 values for as long as both platforms have been measured, and the bound was hiding that behind a magnitude.

A second record is now a maintenance obligation. Adding a scene to the roster means recording it on every platform that has an entry, and a scene added on darwin alone will fail on linux naming itself — which is the intended behaviour and is also more work than before. --fill-baseline covers the running platform only, by construction.

This is what made CI red. With the pnpm double-declaration resolved in 5252bb9 and the themed exit path already distinguishing a declared must-fail from a broken expectation, this bound was the last thing standing between the repository and a green run. Run 33528483192 printed 21 FAIL lines, every one of them a declared must-fail scene failing its declared check, and zero BROKEN lines; its exit code came from these two Scope values and nothing else.