# W14 Worked Case — Six Inert Abuse Cases and a Release Stop

## Case status and safety boundary

This case uses only the frozen L07 fixture. It is offline, deterministic, course-owned, and synthetic. No model, browser, network, external endpoint, account, credential, personal record, exploit, executable prompt, payload, evasion sequence, scan, load generator, or third-party service is involved. The case names defensive scenario classes and controls without explaining how to perform a live attack.

The evidence routes have strict roles. PAPER-37 and PAPER-40 are core bounded research routes. PAPER-07, PAPER-11, PAPER-13, PAPER-34, PAPER-39, and PAPER-41 extend threat, dynamics, mechanism, recommendation, citation, and detection questions. G08 is a candidate tooling reference only. S01 and S02 are voluntary risk-management candidates, not law, certification, safe harbor, or proof of implementation.

## 1. Defensive question and authorization

**Question:** Given six inert scenario records and one declared control per case, what can fail, who can be harmed, which evidence supports a control outcome, and what blocks release?

The authorization is frozen in `system_model.json`: local fixture ingestion, retrieval, annotation, and release review using course-owned synthetic files, with no external system or user traffic. Excluded objects are live answer engines, credentials and accounts, real personal data, and executable adversarial payloads. The learner may read files, run the local standard-library script, compare results, and draft a decision. The learner may not alter a live source, contact a service, exercise a model attack, or persist harmful material.

This scope is proportionate because the learning questions concern joins, control coverage, status vocabulary, risk scoring, decision rules, evidence preservation, and rollback fields. Static CSV/JSON plus a deterministic script provide the necessary evidence. A live test would add risk without answering a necessary course question.

Stop on target ambiguity, personal data, external connectivity, real credentials, unreviewed harmful detail, unexpected output persistence, status uncertainty, inaccessible evidence, or failed rollback. If an input hash changes, stop and inspect version control rather than normalizing the difference.

## 2. System, assets, actors, and trust boundaries

The frozen system is `SYS-RISK-DEMO-001`, an offline synthetic corpus-to-evidence teaching pipeline. Components are ingest, normalize, retrieve, annotate, and release.

| Asset | Protected property | Affected parties |
|---|---|---|
| ASSET-01 | source and citation integrity | reader, source creator |
| ASSET-02 | claim-evidence lineage | researcher, reviewer |
| ASSET-03 | privacy and authorization | data subject, operator |
| ASSET-04 | accessible evidence presentation | assistive-technology user, reader |
| ASSET-05 | governance-status accuracy | decision maker, operator |

Actors are ACT-01, the authorized learner; ACT-02, the independent reviewer; and ACT-03, the untrusted corpus contributor. Capabilities are intentionally narrow. ACT-01 may run and inspect the local fixture. ACT-02 may review evidence and the course decision. ACT-03 exists only as the abstract source of an inert record. No actor has network, credential, live-write, or payload capability.

TB-01 separates the untrusted corpus from ingest. TB-02 separates analysis outputs from release. The entry points are ingest, normalize, retrieve, annotate, and release. This topology matters: source identity, dependencies, instruction/content separation, and synthetic-only privacy controls operate around TB-01; accessibility and governance-status checks operate around TB-02.

## 3. Six scenario-class records

The scenarios are abstractions, not procedures:

| Case | Scenario class | Asset | Entry/boundary | Plausible harm |
|---|---|---|---|---|
| AB-01 | unresolved source identity | ASSET-01 | ingest / TB-01 | a claim edge cannot be traced to a stable source |
| AB-02 | dependency and duplicate amplification | ASSET-02 | retrieve / TB-01 | related records appear to be independent agreement |
| AB-03 | instruction/content boundary failure | ASSET-02 | normalize / TB-01 | untrusted content is treated as operator control |
| AB-04 | privacy exposure | ASSET-03 | ingest / TB-01 | an identifier reaches unauthorized processing or release |
| AB-05 | accessibility and attribution loss | ASSET-04 | release / TB-02 | an assistive-technology user loses the claim-source relation |
| AB-06 | governance-force collapse | ASSET-05 | release / TB-02 | voluntary guidance is represented as mandatory law |

The scenario table is sufficient for selecting defenses. It does not contain a payload, attack string, target, or bypass technique. Any request to add such material exceeds the package boundary.

## 4. Control lifecycle trace

Each case has prevention, detection, response, stop, rollback, owner, and a threshold.

**AB-01 / CTRL-01.** Prevention is a stable-ID allowlist. A failed source join is the detection signal. The unresolved record is quarantined. Any unresolved public claim edge stops release. The rollback restores the last registry snapshot and rebuilds the graph. Owner: data steward. Frozen result: blocked, leaving low/rare residual risk.

**AB-02 / CTRL-02.** Prevention records canonical identity and dependency. Detection flags high overlap or shared origin. Response collapses dependent evidence and rescores. An unreviewed cluster stops an aggregate-evidence claim. Rollback restores the pre-dedup audit view while retaining both views. Owner: retrieval lead. Frozen result: detected, leaving medium/possible residual risk.

**AB-03 / CTRL-03.** Prevention separates inert source content from operator configuration. Detection treats instruction-like markers as untrusted text. Response quarantines the record and reviews the parser boundary. Any propagated source instruction stops the test. Rollback restores the clean corpus and parser configuration. Owner: safety owner. Frozen result: blocked, leaving medium/rare residual risk.

**AB-04 / CTRL-04.** Prevention uses a synthetic-only fixture and field allowlist. Detection scans for identifier-shaped fields before release. Response quarantines the file and notifies the privacy owner. Any real personal data stops and blocks release. Rollback restores an authorized redacted snapshot. Frozen result: blocked, leaving medium/rare residual risk.

**AB-05 / CTRL-05.** Prevention requires semantic claim-source linkage and an accessible table alternative. Detection combines structural checks with keyboard and assistive-technology relation review. Response blocks release and repairs markup. Any lost relation stops release. Rollback restores the last accessible artifact. Owner: accessibility owner. Frozen result: escaped, leaving high/likely residual risk.

**AB-06 / CTRL-06.** Prevention uses controlled status, date, jurisdiction, applicability, and ceiling fields. Detection rejects missing or collapsed labels. Response corrects the crosswalk and requests qualified review. Uncertain force stops the governance claim. Rollback restores the last verified crosswalk. Owner: governance owner. Frozen result: detected, leaving medium/possible residual risk.

Detection after exposure is not prevention. The two detected cases remain visible. An escape is not erased because the overall script passes.

## 5. Detector card and evidence ceiling

Suppose a reviewer proposes a generic detector for instruction-like or optimization-style content. The safe detector card records the intended family, domain, language, modality, observation point, data version, positive and negative denominators, ground-truth process, threshold selection, held-out family, repetitions, uncertainty, false-positive rate, recall, worst-family result, review path, and abstention rule.

The card cannot claim universal coverage. PAPER-37 reports strong performance for some known templates and weak performance for several new black-box families. Its differing false-positive cells show why denominators belong with rates. PAPER-41 reports a strong in-domain detector result but explicitly cannot make a detected style equivalent to falsehood, malicious intent, or policy violation. The appropriate action for a detector alert is controlled review, not automatic accusation.

PAPER-40 adds a different caution: aggregation across contaminated and dependent evidence can fail in condition-specific ways. No universal aggregator winner follows. PAPER-07 and PAPER-13 motivate bounded adversarial-search and citation-integrity questions. PAPER-34 and PAPER-11 motivate bounded mechanism and dynamics questions. PAPER-39 motivates hard-constraint and recommendation-risk review. None authorizes live testing.

## 6. Five-way triage record

Use the five dispositions as a sequence, not a single score:

| Case | Immediate disposition | Required next action | Disclosure assessment |
|---|---|---|---|
| AB-01 | sandbox, then mitigate | retain quarantine, resolve stable identity, rebuild and verify graph | internal data-steward notification |
| AB-02 | sandbox, then mitigate | adjudicate dependency cluster and recompute evidence without synthetic consensus | internal retrieval-owner notification |
| AB-03 | sandbox; stop if propagation occurs | preserve boundary evidence and validate restored corpus/configuration | notify safety owner; external disclosure not indicated by fixture |
| AB-04 | stop on any real data; synthetic case remains sandboxed | quarantine, minimize preserved evidence, validate authorized redacted state | privacy-owner assessment determines any further duty |
| AB-05 | stop external release and mitigate | preserve escape, repair accessible relation, retest manually and programmatically | notify accessibility and release owners; assess affected-party communication |
| AB-06 | mitigate; stop governance claim until verified | correct status and obtain qualified applicability review | notify governance owner; correct any released misstatement if applicable |

No case currently supports package-level **allow** for external release because AB-05 blocks the release. Allow can be reconsidered only after mitigation, regression testing, recovery validation, residual-risk review, and accountable approval. Sandbox remains authorized for the inert fixture.

Disclose is not synonymous with publish. It initiates a scoped assessment: who needs to know, under which policy or requirement, what evidence can be shared safely, and who approves the communication. A public accusation based on a detector score would be unsupported.

## 7. Deterministic L07 reproduction

From the workspace root:

```bash
out_dir=$(mktemp -d /tmp/w14-l07.XXXXXX)
python3 course-labs/L07_drift_adversarial_governance/scripts/review_risk.py \
  --system course-labs/L07_drift_adversarial_governance/data/system_model.json \
  --cases course-labs/L07_drift_adversarial_governance/data/abuse_cases.csv \
  --controls course-labs/L07_drift_adversarial_governance/data/controls.csv \
  --results course-labs/L07_drift_adversarial_governance/data/test_results.csv \
  --governance course-labs/L07_drift_adversarial_governance/data/governance_crosswalk.csv \
  --vocabulary course-labs/shared/fixtures/governance_risk_vocabulary.json \
  --output "$out_dir"
```

Expected stdout is `L07 PASS: 6 cases, 1 escape(s), release BLOCKED`. The parsed `release_decision.json` must show audit status `PASS`, release decision `BLOCKED`, blocking case `AB-05`, outcome counts blocked 3, detected 2, escaped 1, and human approval required. The claim ceiling is synthetic control tests and status-vocabulary audit only.

The W14 validator checks the six frozen input hashes and all eight deterministic output hashes. It uses a fresh temporary directory, checks the file set, parses the decision, and fails closed on drift.

## 8. Evidence preservation and notification

Preserve the six input identities, command, standard output, generated manifest, decision, case results, control coverage, governance audit, threat model, memo, rollback playbook, reviewer observations, and deviation log. Hashes establish file identity, not truth or legal status. Do not place real personal data or harmful detail in the course record.

For AB-05, preserve both the automated structure result and manual relation-review observation. This is important because retaining only the automated PASS would conceal the escape. Notify the accessibility owner and release owner. If an external user or release existed, qualified review would determine any affected-party or public communication; the synthetic fixture does not create an external incident.

Recovery requires an accessible claim-source relation that is programmatically and manually perceivable, regression tests on neighboring cases, restored dependent artifacts, and a recorded residual-risk decision. Rollback availability is not recovery evidence until the protected asset is verified.

## 9. Candidate and standards status

G08 can support a future local test-matrix design, but no G08 execution or generated score is needed here. S01 can organize Govern, Map, Measure, and Manage reasoning. S02 can organize generative-AI risk and action categories. Both standards candidates are voluntary references whose applicability and revision status require refresh. Neither makes the fixture compliant, certified, safe, or legally adequate.

The course policy controls this exercise. It is not authorization to test someone else’s system. External legal, contractual, regulatory, and standards applicability remains outside the package and requires qualified review.

## 10. Release decision and learner deliverable

The correct decision is **STOP external release**. L07’s audit process passes, but AB-05 escapes with high/likely residual risk and meets the blocking rule. The team must sandbox only the frozen fixture, mitigate the accessible relation, preserve evidence, notify named internal owners, validate recovery, and record residual risk. Disclosure beyond those owners requires a separate scoped assessment.

The learner submits a versioned threat model, capability/authorization card, proportionality rationale, six-row control lifecycle, detector card with coverage and false-positive limits, five-way triage table, L07 reproduction record, evidence-preservation index, notification/recovery plan, residual-risk register, and bounded release memo.

A defensible conclusion is: “Within the authorized offline L07 fixture, six inert scenario classes traverse two trust boundaries and produce three blocked, two detected, and one escaped outcomes; detector evidence remains family- and denominator-bound, AB-05 requires stop and mitigation, audit PASS does not approve release, and no live attack, compliance, certification, legal, or production-safety claim follows.”
