# W16 Worked Case — Defending and revising `CAPSTONE-DEMO-001`

## Case status

This is an authored teaching audit of the synthetic **L08 Capstone evidence release**. It is not an observed learner defense. The bundled fixture is course-owned, redistributable, contains no personal data, and uses twelve deterministic synthetic events. The W16 local run reproduces its arithmetic and machine audit. No actual eight-minute presentation, seven-minute defense, forty-eight-hour response, independent reviewer reproduction, accessibility review, pilot, grading calibration, external review, or recording is claimed.

The case asks whether a release candidate supports the sentence “the intervention improves production GEO visibility.” It does not. The audit replaces that sentence with a bounded descriptive claim, preserves an adverse secondary finding, keeps the rejected sentence in provenance, and demonstrates how reviewer findings propagate into a new release identity.

## 1. Freeze the decision before opening the dossier

**Question.** In the fixed synthetic fixture, how do the outcome and citation-correctness proportions differ between treatment and control?

**Estimands.** Primary: treatment-minus-control difference in the fraction of six events per arm with `outcome=1`. Secondary: the same difference for `citation_correct=1`.

**Audience.** W16 learners and authorized reviewers studying evidence-release mechanics.

**Decision.** Route the dossier to human review only if all required files, hashes, artifact references, figure metadata, and frozen calculations agree. Block or weaken any claim with a broken edge, overstated scope, failed material gate, or unreproduced required result.

**Claim ceiling.** Descriptive arithmetic for a deterministic twelve-event synthetic fixture only. The dossier does not estimate production GEO visibility, causal impact, ranking, citation gain, user value, or external validity.

This freeze makes the audit falsifiable. The author is not allowed to redefine “outcome” after seeing the numbers or to treat a structural `PASS` as scientific acceptance.

## 2. Reproduce the twelve-event arithmetic

The source table has six control rows E1–E6 and six treatment rows E7–E12. Manual counting gives:

| Measure | Control positives / total | Treatment positives / total | Treatment − control |
|---|---:|---:|---:|
| Outcome | 2/6 | 4/6 | 0.3333333333333333 |
| Citation correctness | 5/6 | 4/6 | −0.1666666666666667 |

For the primary measure, control rate is (2/6=0.3333333333333333); treatment rate is (4/6=0.6666666666666666); the difference is (2/6=0.3333333333333333). For the secondary measure, control rate is (5/6=0.8333333333333334); treatment rate is (4/6=0.6666666666666666); the displayed difference is approximately (-0.1666666666666667).

The standard-library `analysis.py` produces `result.json`. The L08 validator semantically compares that output with `expected_result.json` rather than relying on filename or raw textual order. A clean run prints:

```text
L08 analysis PASS: 12 events, 2 metrics
L08 PASS: 3 claims, 16 artifacts, analysis_reproduced=true, decision READY_FOR_HUMAN_REVIEW
```

The audit then reports three claims, sixteen artifacts, twelve claim–artifact edges, one public headline, one figure, nineteen required files, eighteen verified checksums, five named human gates, and zero machine errors. These are integrity facts, not a scientific score.

## 3. Audit `CAP-001`, the bounded headline

`CAP-001` states that the synthetic treatment arm has four outcomes among six events versus two among six control events, for a descriptive difference of one third. Its artifact path is:

> `CAP-001` → `DATA-001` → `CODE-001` → `RESULT-001` → `FIG-001` → `LIMIT-001`

The chain is shown linearly for reading, but the ledger stores five direct claim-to-artifact edges. `DATA-001` locates `observations.csv`; `CODE-001` identifies `analysis.py`; `RESULT-001` identifies `expected_result.json`; `FIG-001` identifies the accessible SVG result representation; and `LIMIT-001` carries the descriptive-only ceiling.

The reviewer checks more than existence. The data manifest must declare a synthetic, redistributable input; code must read the declared columns and reject malformed rows; result values must match manual arithmetic; figure metadata must name sources `DATA-001` and `RESULT-001`; the SVG must contain a title, description, and image role; and the limitation must travel with the public sentence.

The result survives these checks. The broad interpretation does not. A reviewer may repeat the descriptive sentence, but cannot infer production improvement.

## 4. Audit `CAP-002`, the adverse secondary result

`CAP-002` records that treatment citation correctness is four in six while control is five in six, giving an adverse treatment-minus-control difference of approximately minus one sixth. Its edges are:

> `CAP-002` → `DATA-001`, `CODE-001`, `RESULT-001`, `NEG-001`, `LIMIT-001`

`NEG-001` makes the adverse direction durable. Without it, a technically complete primary release could still invite a selectively favorable narrative. The reviewer requests that `CAP-002` appear next to the primary result in the research brief and presentation, not only in a deep appendix.

The finding is not described as statistically significant. The fixture supplies neither an uncertainty model nor a population sampling argument. “Adverse” means only that the sign runs against the direction a promoter might prefer. Plausible explanations—noise, measurement tension, or a real trade-off—remain unresolved. The correct response is visibility plus restraint.

## 5. Retain `CAP-003` as rejected provenance

`CAP-003` is the sentence “the intervention improves production GEO visibility.” Its status is rejected. Its only edges are to `NEG-001` and `LIMIT-001`, because no production source, visibility outcome, causal design, or external-validity artifact exists.

The reviewer refuses three proposed repairs. First, renaming the primary chart “visibility gain” would relabel the measure without adding evidence. Second, adding confidence language would not create a production sampling frame. Third, citing the successful validator would confuse package integrity with effect validity.

The revision therefore retains `CAP-003` in the ledger with its rejection reason and removes it from the public headline. This is not an embarrassing leftover. It is evidence that the claim ceiling operated when rhetorical incentives favored a broader story.

## 6. Inspect deviations, exclusions, and reproduction history

The protocol-deviation record states that no row, outcome definition, or analysis changed after the plan. The validator was added later as a packaging control; the estimand remained unchanged. No exclusions are present in the twelve-row fixture. The validator checks all rows and expected artifact paths rather than accepting an unexplained subset.

For the worked audit, the reviewer records one unsuccessful invocation from the wrong working directory. The command could not resolve the dossier path. The error, diagnosis, corrected command, and successful rerun are preserved in `review-findings.md`; the failure does not change the result but reveals why reviewer instructions must use paths relative to a declared release root.

If the corrected clean run had produced different arithmetic, the release would have stopped. The team would preserve both outputs, diagnose input identity and environment, and weaken or remove `CAP-001` until reproduction succeeded or the failure was fully bounded. A final success must not erase a meaningful failed attempt.

## 7. Verify the release identity and five review outputs

The release validator creates exactly:

1. `release_audit.json` — structured counts, status, decision, ceiling, errors, and human gates;
2. `claim_traceability.csv` — the twelve declared claim–artifact edges;
3. `computed_checksums.sha256` — reviewer-computed identities;
4. `reviewer_instructions.md` — commands and interpretation boundaries; and
5. `run_manifest.json` — execution provenance.

Input identities are frozen. Key SHA-256 values include `a95e5f0e5762545396b22e343d98b8f17470028afd2a74780d6d194781dc1d0d` for the L08 validator, `f896e0b49c90937f72e0a28b83193ada1ab22fd139311d1e15bb198510a22e07` for `analysis.py`, `e2545e4f23a3df5171a7e0c24bd6afe80ad58f413d2044df39587bf9a53b335c` for the observations, and `0b1fad6b5611cd092ee183fadb23bd57031edbd5c230b9daebd2a6da861d1171` for the claim ledger.

A checksum proves byte identity, not truth, authorization, accessibility, or fitness. Eighteen verified file hashes do not average away a broken human gate.

## 8. Demonstrate a non-redistributable dependency boundary

The actual L08 fixture is fully redistributable. To teach the alternate route without adding restricted material, the reviewer creates a hypothetical record `DEPENDENCY-NR-01`. It contains no source bytes. The card requires issuer, collection title, edition or retrieval date, exact locator, lawful acquisition instructions, license/access conditions, privacy and security class, local SHA-256, authorized reviewer role, verification method, last successful retrieval, and unresolved scope.

Three states are allowed:

- **Independent reproduction:** an independent authorized reviewer lawfully acquires the exact version, matches the checksum, and regenerates the primary result.
- **Controlled output verification:** the reviewer can inspect identity and selected outputs inside an approved environment but cannot rerun the full chain.
- **Unreproduced:** access or identity cannot be verified; the affected claim is weakened or blocked.

The team never uploads restricted bytes, credentials, cookies, screenshots of protected content, or account state. A hash does not authorize possession. If instructions are too vague for an authorized reviewer to obtain the exact version, the boundary fails verification.

## 9. Apply the five hard release gates

The gate card reads:

| Gate | Fixture evidence | Worked-case status |
|---|---|---|
| Authorization | Course-owned synthetic dossier; declared reviewer role | Pending named human confirmation |
| Privacy | Manifest states no personal data | Pending human privacy review |
| License | Redistributable fixture with recorded license route | Pending human license review |
| Safety | Inert arithmetic; no live system or credential access | Pending contextual safety review |
| Accessibility | Figure metadata, SVG title/description, alt text | Pending human document/figure review |

All five are conjunctive. The automated release decision remains `READY_FOR_HUMAN_REVIEW`, not public approval. An unresolved gate prevents external release even if every checksum matches.

## 10. Conduct the authored defense challenge

The planned eight-minute presentation uses one minute for question/decision, two for design and authorized evidence, two for primary plus adverse results, two for traceability and gates, and one for the bounded conclusion and strongest limitation. It is followed by a planned seven-minute defense. These are course budgets, not observed timings.

The teaching review produces four findings:

| Finding | Challenge | Decision | Evidence-led change |
|---|---|---|---|
| R1 | Headline implies causal production visibility | Accept | Remove that wording; retain rejected `CAP-003`; use bounded `CAP-001` |
| R2 | Adverse secondary result is visually subordinate | Accept | Elevate `CAP-002` and link `NEG-001` beside primary result |
| R3 | Narrative changes invalidate old identities | Accept | Regenerate dependent artifacts, checksums, run manifest, and release notes |
| R4 | Machine-readable alt text is called accessible | Accept | Replace claim with “authoring controls present”; keep human accessibility gate pending |

No row records a completed live defense. It is a model response showing how challenges link to stable coordinates.

## 11. Draft the forty-eight-hour response and revised release

The authored response-to-reviewers states:

> **R1 accepted.** Production and causal wording was unsupported by the outcome, design, and setting. The headline now repeats `CAP-001` within `LIMIT-001`. `CAP-003` remains rejected in the ledger.  
> **R2 accepted.** `CAP-002` now appears at equal text prominence with the primary descriptive difference and routes to `NEG-001`. No significance claim was added.  
> **R3 accepted.** The candidate is versioned as a revision; analysis, figure dependencies, trace export, manifests, checksums, and release notes are regenerated before review.  
> **R4 accepted.** Structural accessibility controls are present, but accessibility conformance is not claimed. Human review remains open and blocks public release.

The response is due within forty-eight hours of a real defense. Here it is a template, not evidence that the clock started or a learner complied. The prior release is preserved. The revised release notes name every changed claim and artifact, whether numeric results changed, which gates were rerun, and what remains pending.

## 12. Final bounded disposition

The machine-readable disposition is `READY_FOR_HUMAN_REVIEW`. The scientific and public disposition is **not yet approved**. `CAP-001` survives only as deterministic descriptive arithmetic. `CAP-002` remains visible as an adverse secondary result. `CAP-003` remains rejected. Authorization, privacy, license, safety, accessibility, the defense, independent review, and external approval still require humans.

Self-check:

- Can every public claim reach valid data, code, result, representation, limitation, reviewer, and release identity routes? Yes for the authored fixture.
- Can another party reproduce it? The local deterministic route is specified and locally rerun, but this package itself is not independent reproduction.
- Are negative results and deviations retained? Yes.
- Is a non-redistributable route precisely bounded without restricted bytes? Yes, as a hypothetical teaching card only.
- Are hard gates complete? No; they are explicitly pending.
- Does local PASS establish human review, accessibility conformance, pilot quality, grading calibration, external review, publication readiness, or recording? No.

The result is worth trusting only to the extent that its narrow claim, artifact chain, dependency boundary, retained counterevidence, and open gates remain visible. Trust here means readiness for accountable judgment, not guaranteed agreement.
