ADR-0111 The site checks its own copy
Extends ADR-0018 and playbook §6 to a corpus they did not cover.
#Context
content-lint has thirteen rules and a scan that runs them over real shipped copy. The corpus was component strings: every component carrying tone metadata, scanned by calling its built tone function with the copy in its own Storybook story. That is the copy a product ships.
It is not the copy most people read. The documentation site authors 229 strings of its own, and the site that explains the voice was the one surface with no check on it. The gap was not an oversight in the rules. It was that the scan knew how to find a component's copy and nothing else.
The prose that accumulated there had a consistent fault, and it is worth naming because the rules cannot catch it. The subject of the sentence was the machinery:
Every cell the generator wrote, whole. A contract with no dimension has no pairing fact to read out of code, and is listed rather than given a written row.
Every figure below is read out of a committed receipt rather than written here: the engine wrote the scene numbers, the generator wrote the contrast margins, and the scene table declared which fixtures are supposed to fail.
A reader wants to know what a number means. These sentences say which internal tool produced it. The second also re-argues the site's honesty, which the footer already states once for every page, and points with "below", which directional-language forbids in component copy and had no way to see here.
#Decision
- **
content-lint:scan-sitereadsapps/docs/app/site/text.tsand runs the
rules over every string it authors.** The module is imported rather than parsed: it declares no imports of its own, so Node's type stripping loads it directly, and a sentence concatenated across three source lines arrives as the one sentence a reader sees. A regex over the file reports that as three fragments, each ending mid-clause.
- Two corpora, because two kinds of string live there. A label names a
thing; running prose explains one. Prose is excused two rules:
spelled-out-numeralwants3, which is right on a label beside a figure
and wrong in a sentence. "The same system, entered from the 3 jobs that use it" is not the house voice, it is a worse sentence.
sentence-casereads two capitals after the first as Title Case. In a
paragraph the second sentence begins with one, so every multi-sentence string trips a rule about headings.
Neither is a defect in those rules. Both were written for short UI strings and are correct there. The discriminator is mechanical, which §6 requires: a string ending in a sentence's own punctuation is prose, everything else is a label.
- Fixed prose moves into
text.ts. That module's own header says it holds
every fixed string the site renders, authored once. Four scene verdicts in body-receipts.tsx and two strings in body-component.tsx were not there, so the check could not see them. They are there now, which is what puts them under it.
- A template is not checked and is not claimed to be clean. A string the
module builds from a number it is handed would have to be called with an invented number to check it. The scan skips those rather than fabricating an argument.
#Consequences
229 strings checked, 0 violations. Run against the copy as it stood before this record, the scan reports exactly one violation: the "below" quoted above.
The two excused rules were not excused to make that number look better. A first pass that applied every rule to every string reported seven more, all of them from those two rules and none of them a defect in the copy: five demands for a digit inside a sentence, and two paragraphs read as headings in Title Case. That pass is what produced the split.
Eleven strings were rewritten off the machinery and onto the subject, and six moved into the checked module, three of those unchanged. No fact was dropped: the sizes, the widths and the counts are all still there.
The scan was confirmed to fail before it was left passing. Three violations planted in text.ts, one prose and one label and one courtesy word, were each reported with their key path, and the tool exits non-zero.
The body modules still hold prose that is built from a contract's own values, and a template's fixed words are still unchecked. Both are stated here rather than implied by a green scan.
#Alternatives
Scan the rendered HTML of all 389 pages. Catches every string wherever it is authored, including the body modules and the templates. Rejected for now because it needs a production build to run at all, and a check that cannot run without one is a check that runs late.
Parse text.ts with a regex. No type-stripping flag. Rejected on the evidence: it reports a concatenated sentence as fragments, and the first probe written that way produced a violation whose quoted text ended in : the.
Apply every rule to both corpora. One rule set, no discriminator. Rejected because it is wrong twice over: it would demand digits in running prose, and it would report every paragraph as a heading in Title Case.