1. Structure is a verification mechanism, not a ranking recipe
Content architecture matters because readers and information systems operate on representations, not on an organization's private fact ledger. A true statement hidden inside a promotional paragraph can be hard to verify, easy to segment away from its qualification, and difficult to update consistently. A well-structured evidence page makes identity, scope, claims, evidence, method, time, limitations, and revision status locally recoverable. That is a defensible design objective even if no ranking or citation metric moves.
The same observation creates a danger: useful editorial practices are easily turned into superstition. A heading, table, definition, or structured-data record is not a universal signal. Its effect depends on the human task, parser, segmentation, corpus, query, system, date, and outcome. An evidence page should therefore be judged first on factual integrity, usability, accessibility, and governance. Any system-response hypothesis is tested separately at the stage the intervention can plausibly reach.
W09 distinguishes four claims:
- A structure can make a claim easier for a human to locate and verify.
- A structure can make a representation more consistent for a declared parser.
- A controlled open pipeline can show a stage-specific retrieval or use contrast.
- A live platform can display a time-bounded outcome under recorded conditions.
Evidence for one does not automatically establish the next. L05 directly checks a portion of the second claim and prepares a later human-comprehension test. It does not collect the third or fourth kind of evidence.
Checkpoint A. A page passes semantic HTML and schema validation. What is established? Only the declared validation properties under the tested tools and versions. Retrieval, ranking, citation, answer fidelity, user action, and commercial effect remain unmeasured.
2. Begin with a governed fact system and atomic claim lock
An evidence page should be generated from governed claims rather than used as the only place facts live. The internal record for each claim needs a stable ID, canonical entity, minimal proposition, value and unit, time interval, locale, applicability conditions, evidence spans, source identities, independence, owner, reviewer, status, contradiction state, and permitted public wording. The page is one view of that system.
A compound sentence can conceal several propositions. Consider: “Our certified service cut review time by 40% for global teams.” At least five questions appear. Which service version? What certification and validity period? What outcome definition? Which comparator and method produced 40%? Which teams, locations, and dates? A structural rewrite must not split the sentence by deleting these conditions. Atomicity means each independently challengeable proposition has an ID and evidence path; it does not mean every sentence must be short.
For a controlled intervention, freeze a claim lock:
where (c_i) is claim ID, (t_i) exact approved text or normalized proposition, (S_i) source identities, (\kappa_i) material conditions, and (v_i) version. The control and treatment are acceptable only if every locked claim and source relation remains equivalent under the chosen contract.
The L05 fixture uses exact text equality for three synthetic claims and set equality for three source IDs. A real project may allow controlled expression-preserving changes, but then factual equivalence needs reviewed propositions, units, quantities, conditions, entailment, and attribution—not a similarity score alone.
Checkpoint B. Two sentences have embedding similarity 0.97, but one omits “among completed responses.” Are they factually equivalent? No. A material denominator condition changed even though a text-similarity proxy is high.
3. Reconstruct an evidence page by function
An evidence-rich page has components with explicit jobs:
| Component | Verification function | Failure made visible |
|---|---|---|
| Identity and version header | names entity, owner, status, locale, and last substantive review | alias confusion and stale page state |
| Answer-first scoped summary | states the main proposition and critical condition early | a conclusion detached from its population |
| Atomic claim blocks | preserve claim ID, proposition, qualification, and evidence edge | one citation laundering adjacent claims |
| Evidence table | binds value, unit, comparator, method, date, and source | incomparable numbers and missing denominators |
| Definitions | disambiguate technical, legal, and commercial terms | query/entity mismatch and scope drift |
| Method | explains collection, transformation, analysis, and review | uninspectable result production |
| Limitations and unknowns | states failure conditions and excluded transfer | a local result presented as general |
| Source identities | resolve origin, version, passage, and dependence | circular or wrong-source attribution |
| Change history | records corrections, supersession, and reason | silent mutation and freshness ambiguity |
| Correction or next-action route | provides accountable update and evidence access | no way to challenge or refresh a claim |
This anatomy is not a template that every page must reproduce in the same visual form. A short specification page and a long research report can express the functions differently. The invariant is recoverability: a reviewer can locate the content needed to assess a claim without inferring missing scope.
One original visual for the lecture is a vertical page reconstruction with a parallel “verification rail.” Each page component connects to the question it answers: what object, what claim, under which conditions, based on what, how produced, where it fails, what changed, and who can correct it. A banner states: “anatomy, not outcome guarantee.”
4. Co-locate claims with scope, method, date, and limitation
Page segmentation can separate a number from the sentence that qualifies it. Co-location reduces this risk. A value should travel with its unit, population, comparator, method, date, uncertainty, and source. A definition should remain near the term. A limitation that materially changes interpretation should not be hidden in a remote footer.
Co-location does not require repetitive keyword stuffing. It means designing semantic units that remain meaningful when excerpted. A compact evidence block might contain:
- claim ID and plain-language proposition;
- status and validity date;
- value, unit, comparator, and population;
- method and evidence locator;
- material limitation or uncertainty;
- source identity and correction owner.
Tables are useful when cells are genuinely comparable. If two metrics have different denominators or dates, either normalize the comparison under an explicit method or keep them separate. A clean-looking table can amplify error when its columns imply equivalence that does not exist.
Dates require types. Publication date, observation window, source-validity interval, last reviewed date, and next review date are different. “Updated recently” is not a version field. The change history should distinguish typo correction, evidence refresh, claim change, method change, and structural intervention.
Limitations are not negative marketing. They are part of the claim. They show a reader where the statement stops and provide counterevidence to an overbroad extraction. Removing a limitation to make a treatment more concise violates factual equivalence.
5. Keep visible content, metadata, and schema consistent
An evidence page can expose several representations: visible HTML, title and description metadata, link text, structured data, feeds, downloadable files, and accessibility text. These representations must not make incompatible claims. Structured data cannot promote an unpublished award, unsupported rating, broader service region, or different price while the visible page remains qualified.
Build a consistency matrix with one row per atomic claim and columns for visible block, metadata, JSON-LD or other schema, source link, alt/caption where applicable, and version. Each cell records equivalent, absent by design, conflicting, stale, or unreviewed. “Absent” can be acceptable because not every representation encodes every claim. “Stronger but machine-readable” is not acceptable.
PLAT-03 documents Google's general structured-data guidelines and the distinction between valid markup and eligibility for supported features. Eligibility does not guarantee display. PLAT-13 documents Schema.org vocabulary semantics. A vocabulary term does not establish that a consumer recognizes it, that a platform uses it in a hidden stage, or that visibility improves. PLAT-02 provides Google-specific guidance emphasizing accuracy, quality, relevance, and user value while warning about scaled abuse; it is policy guidance, not an experimentally validated content score.
Schema validation checks syntax and some declared constraints. Factual validation checks whether the encoded proposition has the same evidence, status, scope, and date as visible content. Both are needed. Neither is a ranking experiment.
Checkpoint C. Visible text says “rated 4.2/5 among 120 verified respondents in 2025,” while schema says aggregate rating 4.8 with no count. The page fails factual representation consistency even if the schema syntax is valid.
6. Define factual equivalence before constructing treatment
Factual equivalence protects the interpretation of a structural intervention. At minimum, compare:
- claim IDs and normalized propositions;
- quantities, units, denominators, uncertainty, and rounding;
- entity IDs, aliases, product/version scope, and locale;
- dates, observation windows, and status;
- material qualifications, exceptions, and limitations;
- source identities, evidence spans, and attribution edges;
- disclosures, sponsorship, and authorship;
- correction and update routes when they bear on trust.
Exact text equality is a strong and simple lock for a layout-only test. It is insufficient when the intervention intentionally reorganizes sentences, but looser contracts require more careful review. An embedding similarity or character-diff percentage cannot establish factual equivalence. A reviewer must verify that every proposition and condition is preserved.
Visual equivalence is neither required nor desired when layout is the factor. Structural equivalence is not required when structure is the factor. The experimental contract names what must remain identical and what is permitted to differ. The complete diff then tests that contract.
Use an independent reviewer where feasible. The L05 checklist requires the reviewer ID to differ from the experimenter ID and includes atomic claim text, numbers and qualifications, source identity, attribution, reading order, and rollback review. This separation does not prove the review was competent or complete; it prevents a trivial same-role pass and makes the review record inspectable.
7. Write an intervention card that can falsify the idea
An intervention card turns “make the page clearer” into a testable, governed change. It should record:
| Field | Required decision |
|---|---|
| Intervention ID and version | stable identity for treatment and results |
| Environment and authorization | local fixture or explicitly owned staging; owner and limits |
| Factor | exactly what changes |
| Hypothesized stage | human comprehension, parsing, retrieval, generation, attribution, or another declared stage |
| Expected signal | observable outcome and direction, without guarantee language |
| Falsification condition | what result would weaken the mechanism hypothesis |
| Non-claims | stages and outcomes not measured |
| Confounds | alternative explanations and accidental changes |
| Validity/accessibility tests | gates required before evaluation or release |
| Stop conditions | events that halt the work |
| Rollback | retained state, instruction, hashes, and acceptance test |
The L05 card declares evidence_layout_proximity, a local_fixture, and human_comprehension as the hypothesized stage. It states that a later approved comprehension task may show fewer claim-to-source matching errors. The word “may” matters: L05 itself performs no user test. It explicitly disclaims crawler, retrieval, ranking, citation, traffic, and business effects.
A mechanism hypothesis needs at least one alternative. Evidence cards could help because claim and source are spatially closer; they could also help because the border attracts attention, or hurt because the grid causes wrapping. Viewport, familiarity, and assistive-technology behavior remain confounds until tested.
8. Isolate one factor with complete local diffs and tests
A one-factor design fails when any undeclared difference could change the outcome or violate a guardrail. For a layout-proximity intervention, claim wording, order, title, language, link targets, source IDs, metadata, and other style rules should remain locked. The complete HTML/CSS diff is evidence; a screenshot comparison is not complete because it can hide DOM, metadata, focus, and off-screen changes.
L05 normalizes the declared stylesheet filename and requires the two HTML files to match after that substitution. It then isolates the .evidence-block CSS rule and requires all other CSS to match. The generated diff exposes the stylesheet switch and the one changed rule. The parser independently extracts claims and source IDs from both pages and compares them with the lock.
The deterministic audit reports:
- three locked claims;
- three source identities;
- six passed independent-review criteria;
- one changed factor,
evidence_layout_proximity; - zero contract errors;
- a claim ceiling of local structural equivalence only.
This proves a narrow software property. It does not prove the treatment improves comprehension. A later human study would need task, allocation, participants, accessibility, outcome, uncertainty, and ethics. A reconstructed pipeline would address other stages. Keep these studies separate.
9. Make accessibility and validity non-compensable release gates
Accessibility is not polish and cannot be traded for a target metric. Both arms need valid language and title, one clear main, one h1, logical heading structure, source-link purpose, unchanged DOM reading order where required, visible keyboard focus, usable reflow, sufficient contrast, and comprehensible text alternatives. Motion, hover, and pointer precision cannot carry core evidence.
Automated checks have a ceiling. The L05 parser can count elements, inspect selected attributes, and compare source IDs. It cannot determine whether link language is understandable in context, whether focus order matches meaning, whether zoomed layout hides content, whether contrast passes in every state, or whether a screen-reader user can efficiently match claims and sources. Those require browser tests, tools, and human review.
Validity gates include claim/source equality, consistent schema, authorized environment, non-deception, no hidden user-specific variation, no fabricated evidence, no cloaking, and no privacy exposure. If any gate fails, stop or roll back even if a downstream metric improves. A favorable citation or ranking observation cannot compensate for factual drift.
The release record must separate “automated invariant passed,” “manual check passed,” “assistive-technology review pending,” and “not applicable.” Do not collapse them into one accessibility score.
A practical release matrix gives every gate five additional fields: test method and version, artifact or observation that contains the result, accountable reviewer, decision threshold, and expiry or retest trigger. For example, a heading-order check may cite a static DOM inspection, while a claim-to-source navigation check may cite a keyboard and screen-reader observation. A contrast result must identify the tested foreground/background pair and state, not merely name a tool. A factual review resolves to claim and evidence IDs. This structure prevents an undated green badge from standing in for inspectable evidence.
Review order matters. Run claim and authorization checks before investing in presentation polish. Run automated structural tests before manual sessions so obvious failures are removed. Then perform browser, keyboard, zoom/reflow, and assistive-technology review on both arms under the same declared viewport and content state. Finally, an independent release reviewer inspects the complete diff, open findings, deviations, and rollback evidence. Passing one arm is insufficient: treatment integrity is comparative.
If a finding is disputed, preserve the observation and reviewer disagreement. Do not force a pass to meet a deadline. Classify the item as blocking, non-blocking with rationale and expiry, or unresolved and therefore release-blocking under the predeclared rule. A waiver is a governed exception with owner, scope, compensating control, and expiry; it is not evidence that the underlying accessibility or validity requirement was met.
10. Use factorial, sequential, or bundle language when one factor is impossible
Real redesigns often need several coupled changes. Adding evidence cards may require markup, headings, spacing, link text, and responsive rules. Four honest options exist.
- Reduce the treatment. Keep only the smallest interpretable factor.
- Factorial design. For two factors (A) and (B), evaluate control, (A), (B), and (A+B), with enough independent units to estimate main effects and interaction.
- Sequential design. Freeze an order, decision rule, and stopping boundary before seeing outcomes. Each stage has a new version and estimand.
- Bundle estimand. Change the full redesign and claim only the effect of the bundle.
A factorial design is not automatically superior. It increases arms, sample and review burden, and interaction complexity. A sequential design can contaminate later stages if earlier results determine unplanned edits. The choice follows the decision question and available independent units.
Never call a second change “cleanup” merely because it was convenient. If the treatment modifies wording, source links, and layout, it is a compound intervention. The deviation log should record discovery, affected artifacts, effect on the estimand, corrective action, and whether the study resets, downgrades to exploratory, or stops.
11. Test rollback and preserve the deviation log
Rollback is a verified restoration process, not the sentence “we can revert.” Retain the exact control HTML and CSS, their hashes, the treatment hashes, replacement instruction, required permissions, and acceptance tests. Execute the rollback in the local environment and verify the restored hashes. Preserve the treatment and audit evidence separately so the change remains inspectable.
L05 records control HTML SHA-256 fff13e70c2218f8ea607f93128f5845f5d2582e6973badd469143c1b1fc18e78 and control CSS SHA-256 69f8bd330f589b84d309843be1e7ad95af19eac04c4e05da4e5409552182383d. The treatment hashes are distinct. The rollback instruction replaces the treatment pair with the retained control pair and then verifies both. A mismatch is a stop condition.
The deviation log includes planned version, observed deviation, detection time, detector, affected claims and files, cause if known, immediate containment, effect on analysis, approval, and disposition. Null results, adverse accessibility behavior, and unanticipated wrapping belong in the same scientific record. Do not silently repair the treatment and keep the original hypothesis test.
12. Read the evidence routes through their stage and validity ceilings
PAPER-12 evaluates rewriting after a product already belongs to a fixed ten-item slate; negative methods show why no tactic is universally positive. PAPER-21 uses controlled candidate contexts and reports utility regressions in several cells; it cannot establish stable preferences of commercial engines. PAPER-23 reconstructs retrieval, reranking, and citation stages and reports that some body-only changes reduce outcomes. Its stage results support experimental design, not production field weights.
PAPER-08 is observational within an audited union of cited English B2B SaaS URLs, without a prospectively defined non-cited population adequate for causal editing claims. PAPER-09 uses synthetic tourism rewrites and a small simulated visibility study; lexical and model-based quality measures do not establish factual preservation. PAPER-28 remains audit-only: its reported percentages, thresholds, weights, platform labels, and similarity measures are not admitted as W09 findings because units, versions, factual preservation, and reproduction remain unresolved.
The platform route is equally bounded. PLAT-02 can inform a responsible-publishing checklist within Google policy scope. PLAT-03 can inform supported structured-data eligibility and policy checks. PLAT-13 can define vocabulary terms. None proves retrieval, ranking, citation, truth, or business outcome.
A defensible W09 conclusion has two parts. First: “The local deterministic audit verified three identical locked claims and three identical source identities across control and treatment, isolated one declared CSS factor, retained six independent-review checks, and produced rollback hashes.” Second: “No human or external system response was measured; accessibility, user benefit, retrieval, ranking, citation, and production transfer require separate evidence.”
The unit succeeds when structure makes evidence easier to inspect without making the claim stronger than its evidence—and when the release process can prove exactly what changed, what did not, and how to return safely.