# W14 Transcript — Planned No-Video Equivalent

## Status and use

This document is an authored text equivalent for a planned 35–45 minute W14 explanation. It is not a transcript of a recording. No W14 video or audio exists, and no W14 seminar, studio, rehearsal, or timed pilot has occurred. The chapter times are authoring budgets, not observed media duration.

All cases are inert, local, synthetic, and non-operational. Do not add a live target, executable prompt, payload, credential, evasion sequence, load pattern, scan procedure, or third-party endpoint. The package teaches defensive reasoning and evidence preservation, not attack execution.

## Planned chapter budget

| Chapter | Focus | Planned minutes |
|---|---|---:|
| 1 | The boundary between optimization and abuse | 4 |
| 2 | System and threat model | 5 |
| 3 | Authorization and proportionality | 4 |
| 4 | Prevention, detection, and detector limits | 5 |
| 5 | Research evidence ceilings | 4 |
| 6 | Safe triage | 4 |
| 7 | Evidence preservation and response | 4 |
| 8 | Notification, rollback, and recovery | 4 |
| 9 | L07 reproduction and decision | 4 |
| 10 | Integrated release sentence | 3 |

The total is a planning estimate.

## Chapter 1 — The boundary between optimization and abuse

Our essential question is: when does optimization become system abuse? It is tempting to answer with intent. An honest creator optimizes; a malicious actor manipulates. That answer fails quickly. A creator can intend to clarify and still publish a misleading comparison. An actor can use technically true sentences to fabricate a false impression of consensus. A researcher can claim a defensive purpose while testing without permission.

Use seven judgments instead. First, truth: are explicit and implied claims supported within scope? Second, provenance: can sources and dependencies be traced? Third, disclosure: are sponsorship, affiliation, and material interests visible where needed? Fourth, system integrity: does content remain evidence, or is it treated as a control instruction? Fifth, competitive fairness: does the intervention impersonate independence, fabricate consensus, or make unsupported claims about another party? Sixth, authorization: who approved precisely what? Seventh, proportionality: is the expected defensive evidence worth the risk, and could a less intrusive method answer the question?

Notice what is absent: visibility improvement. An evidence-preserving edit is not defined by whether it ranks better. An abusive test is not legitimized because it produces interesting measurements.

This course also has a hard negative boundary. We do not conduct live poisoning, evasion, credential use, unauthorized load, harmful payload release, cloaking, deceptive attribution, or third-party testing. We use inert scenario classes. A class such as “dependency amplification” is enough to identify an asset, harm, and control. We do not need a recipe for causing it.

[Pause. Imagine an edit whose claims are correct but whose pages all derive from one undisclosed source. Which gates can fail? Provenance, disclosure, and possibly competitive fairness. Correct sentences do not complete the audit.]

The conclusion at this stage is that classification is governed and evidence based. Name the intervention, system version, scope, actors, authorization, affected parties, and evidence. If those fields are unavailable, do not substitute a moral label or a detector score.

## Chapter 2 — System and threat model

Threat modeling begins before threats. Freeze a versioned system. Record its scope, components, data flows, assets, actors, affected parties, capabilities, entry points, trust boundaries, exclusions, and owners.

The L07 system is an offline synthetic corpus-to-evidence teaching pipeline. Its components are ingest, normalize, retrieve, annotate, and release. It contains no live answer engine, credential, account, real personal data, or executable adversarial payload.

Five protected assets make the model concrete. ASSET-01 is source and citation integrity, affecting readers and source creators. ASSET-02 is claim-evidence lineage, affecting researchers and reviewers. ASSET-03 is privacy and authorization, affecting data subjects and operators. ASSET-04 is accessible evidence presentation, affecting assistive-technology users and readers. ASSET-05 is governance-status accuracy, affecting decision makers and operators.

Actors are not unlimited capabilities. ACT-01, the authorized learner, can run the local script and inspect frozen outputs. ACT-02, the independent reviewer, can review evidence. ACT-03, the untrusted corpus contributor, is represented only as the abstract origin of an inert record. None has network, credential, live-write, or payload permission.

Write capabilities as verb, object, scope, time, and approval. “Corpus access” is vague. “Submit one course-owned synthetic record to the inert fixture during the approved exercise” is inspectable. This grammar prevents a local read from silently becoming an external write.

Two trust boundaries organize the cases. TB-01 is the transition from untrusted corpus to ingest. TB-02 is the transition from analysis output to release. Entry points include ingest, normalize, retrieve, annotate, and release.

Now map scenario to harm. An unresolved source identity can break attribution. A dependency cluster can inflate apparent consensus. An instruction/content boundary failure can let untrusted data influence configuration. A personal identifier can violate privacy. An inaccessible relation can prevent a user from associating a claim with its source. A collapsed status label can turn voluntary guidance into apparent law.

[Pause. Choose AB-05. State asset, affected party, entry point, trust boundary, harm, and owner. If you start describing how to break accessibility, stop. The fixture already supplies the failure state; your job is control and recovery.]

Version the model. If a component, actor, permission, or release destination changes, the threat model must be reviewed. A polished diagram is not evidence of current coverage.

Audit the model with completeness and dependency questions. Does every scenario name an asset and affected party? Does every trust-boundary crossing have a prevention or detection opportunity? Does every control have an owner, threshold, response, stop, and rollback? Are two controls dependent on one registry or judge? Which component is outside visibility? Which harms occur even when the primary output appears correct?

Then distinguish exposure from impact. A duplicate record entering a corpus is an exposure. Counting related records as independent support is an integrity impact. A detector flag is an observation. Withholding the aggregate is a containment decision. Conflating those stages makes it impossible to tell whether prevention failed, detection succeeded, or recovery happened.

Also distinguish system owner from affected party. The operator may control release, but an assistive-technology user experiences AB-05. A source creator can be harmed by wrong attribution even if the operator sees no immediate cost. Threat models that list only infrastructure assets can miss research-integrity and accessibility harms.

Finally, record exclusions as claims about scope. “No live systems” does not mean live systems are safe; it means they are not evaluated. “No personal data” means privacy behavior on real data is not tested. Exclusions prevent accidental generalization and guide the next evidence request.

## Chapter 3 — Authorization and proportionality

Authorization is not the phrase “for research.” It is a record. Name the accountable owner, system and asset versions, allowed actors, exact capabilities, approved scenario classes, dates, locations, data classes, methods, retention plan, notification route, stop conditions, rollback authority, and exclusions.

L07’s authorization is narrow: course-owned synthetic files, local processing, and no external user traffic. That record gives the learner permission to reproduce the fixture. It does not authorize a test of a commercial system. The course policy is an instructional baseline, not someone else’s permission.

Authorization can expire or become inapplicable. A new target, model, data source, network, payload class, persistence method, or action capability requires review. A tool can make an operation possible without making it permitted. Candidate G08 may support the design of test matrices, assertions, and CI checks. We do not install or connect it here. Even if we did, its scores would depend on the validity of fixtures, assertions, judges, thresholds, and denominators.

Proportionality selects the least intrusive method that can answer a defined defensive question. List expected evidence, maximum plausible harm, affected parties, reversibility, uncertainty reduction, and safer alternatives. Static records can answer many questions: do case-control joins resolve? Are lifecycle fields complete? Does a release rule block an escaped high-risk case? Does a governance vocabulary preserve status? None requires a live model.

If a static review is insufficient, consider a disconnected simulation. If a disconnected simulation is insufficient, retain the evidence gap or seek a separately governed process. Do not jump to a live action because it seems more realistic.

[Pause. The question is whether a rollback playbook contains an owner and success criterion. What is the proportionate method? Read and validate the frozen record. A live incident would add no necessary evidence.]

Document stops. Missing authorization, target ambiguity, personal data, unknown persistence, absent logs, inaccessible evidence, or untested rollback ends the current route. A stop is a valid outcome and should remain in the research record.

## Chapter 4 — Prevention, detection, and detector limits

Prevention acts before the harmful transition. L07 includes stable source IDs, dependency registration, separation of inert content from configuration, synthetic-only field allowlists, accessible claim-source linkage, and controlled governance vocabularies. The evidence must show a gate ran on the current version. A policy statement is not implementation evidence.

Detection observes a signal during or after exposure. Keep the timing honest. AB-02 is detected and the aggregate count is withheld. AB-06 is detected and the status label is corrected. Neither outcome should be rewritten as prevention.

A detector card names the threat family, domain, language, modality, observation point, model or parser version, sample route, ground truth, positive and negative counts, threshold selection, held-out design, repetitions, uncertainty, true positives, false negatives, false positives, true negatives, recall, false-positive rate, worst-family result, and review path.

Coverage is a vector. A detector can perform strongly on known string templates and fail on unfamiliar reasoning or review patterns. A text detector says nothing about images unless evaluated there. An in-domain result says little about a shifted domain without evidence.

False-positive rate requires a denominator. In a low-prevalence population, even a small rate can produce many false alerts. The consequence matters too. Routing an item to private review is not equivalent to public accusation, suppression, or permanent punishment. High-impact responses require stronger evidence, independent review, correction or appeal, and accountable approval.

PAPER-37 is core because the reported evidence makes family coverage visible. Some known-template results are high. Several new black-box-family results are very weak. Reported false-positive evidence also varies by table and sample. The source does not support a universal defense claim.

PAPER-41 extends this lesson to detecting GEO-optimized web-content style. Its benchmark result belongs to the dataset and pipeline. A detected style is not proof that content is false, malicious, illegal, or policy violating. It should route to review.

[Pause. Complete: “This detector has recall ___ on ___ positive cases from ___ family, FPR ___ on ___ negative cases, and unknown coverage for ___.” If you cannot fill those blanks, do not present one score as a release control.]

Consider a fictional detector panel to make the arithmetic concrete. Suppose the frozen evaluation contains 100 positive scenario records and 900 negative records. The detector flags 80 positives and 18 negatives. Recall is 80 divided by 100, or 0.80. False-positive rate is 18 divided by 900, or 0.02. There are 98 alerts, of which 80 are positive under this constructed ground truth. That alert precision is about 0.816 inside this particular panel.

Now change only the population base rate. If deployment contains far fewer true positives, the same conditional recall and false-positive rate can yield a much lower share of true alerts. This is why an evaluation panel’s class balance cannot be treated as prevalence. It is also why a two-percent false-positive rate may be tolerable for low-cost private review but unacceptable for an automated public accusation.

Break the fictional 100 positives into families. If all 80 detected positives come from two familiar families and a third family has zero recall, the overall 0.80 hides the failure that matters. Report the family matrix and uncertainty. If the threshold was selected on these same 1,000 records, the estimates are optimistic for model selection; reserve independent data or use a documented validation/test split.

This arithmetic is not a claim about PAPER-37 or PAPER-41. It is a synthetic teaching example. The papers motivate the reporting questions, while their own values remain tied to their reported samples and designs.

## Chapter 5 — Research evidence ceilings

PAPER-40 is the second core route. It evaluates fact-verification evidence aggregation under constructed GEO-style poisoning. Reported performance degrades in some contamination conditions, and no aggregator wins across every condition. The course uses this to motivate per-condition matrices, source-family dependence, repeated runs, abstention, and uncertainty.

The ceiling matters. Repetition, data dates, annotation agreement, implementation detail, and uncertainty are not sufficient for a universal robustness ranking. We do not reproduce poisoning mechanics. We study the defensive consequence: evidence items that appear separate may share an origin, and aggregation can fail differently across conditions.

PAPER-07 supplies bounded adversarial-search evidence. It does not establish that every source reaches retrieval or that current production behavior is stable. PAPER-13 supplies bounded citation-vulnerability evidence from a time-sensitive setting. It does not establish a universal citation mechanism.

PAPER-34 supplies a bounded mechanistic view of persuasion-related pathways in controlled tasks and open models. It does not prove a complete explanation of free-form or proprietary systems. PAPER-11 supplies a theoretical dynamics route. Its local assumptions can yield non-monotonic behavior, which motivates parameter sensitivity but not a policy proof.

PAPER-39 supplies a bounded recommendation-agent risk and defense benchmark. Its rates belong to the evaluated suite, agents, variants, models, and single-turn conditions. They do not estimate universal prevalence or guarantee a defense. PAPER-41 supplies detection evidence with the accusation boundary already discussed.

G08 is a tooling candidate. S01 is the voluntary NIST AI Risk Management Framework 1.0. S02 is the voluntary NIST Generative AI Profile. These candidates help structure tests and governance. They are not laws, certifications, conformity assessments, safe harbors, or proof of implementation.

[Pause. Pick one route. State the object it evaluates, the setting, the defensive question it supports, and one forbidden extrapolation. Do not summarize a source as simply “proves risk” or “solves defense.”]

The full controlled route is therefore PAPER-37, PAPER-40, PAPER-07, PAPER-11, PAPER-13, PAPER-34, PAPER-39, PAPER-41, G08, S01, and S02. The list is a role matrix, not an authority stack.

## Chapter 6 — Safe triage

W14 uses five dispositions: allow, sandbox, mitigate, stop, and disclose.

Allow means the specified activity may continue under current scope. It does not mean zero risk. Record the residual risk, accepting owner, monitoring trigger, review date, and revocation conditions.

Sandbox means the defensive question may be tested only in an isolated, synthetic environment. Record isolation, input provenance, output controls, logs, cleanup, and rollback. “Sandbox” describes the environment, not an inherently safe scenario.

Mitigate means repair a control or evidence gap, then retest. A proposed fix is not completed mitigation. Preserve before and after artifacts and run regression checks because a change can shift another control.

Stop means halt a specified transition, test, or release. Name what stops, why, who owns the decision, and what evidence could permit resumption. Missing authorization, personal data, inaccessible core evidence, broken logs, or failed rollback can trigger it.

Disclose means initiate a scoped notification or disclosure assessment. It can mean internal owner notification, coordinated vendor reporting, affected-party communication, regulator notice where applicable, or public correction. Those have different evidence and authority requirements. A detector score does not automatically justify public disclosure.

Dispositions can form a sequence. A case begins in the synthetic sandbox, exposes a control gap, triggers mitigation, stops release, and notifies owners. This is exactly how AB-05 should be read.

[Pause. Triage AB-04. The frozen identifier is synthetic, so the case can remain sandboxed. If any real personal data appears, stop, quarantine, minimize preservation, notify the privacy owner, and validate a redacted recovery state. Public disclosure remains a qualified decision.]

The package-level allow state is unavailable because AB-05 blocks release. This prevents a team from declaring general eligibility based on the other five outcomes.

Walk the six cases through the dispositions. AB-01 begins in the sandbox, is blocked by source identity control, then needs mitigation of the record before any graph rebuild. The data steward receives an internal notification. AB-02 remains a sandboxed detection case; the retrieval owner adjudicates the dependency family and recomputes the aggregate before a claim proceeds.

AB-03 is sandboxed only while the marker remains inert data. Any propagation into control state triggers stop. The safety owner preserves boundary evidence and verifies the clean configuration. AB-04 is a synthetic privacy case in the sandbox; any real identifier would trigger immediate stop, quarantine, evidence minimization, and privacy-owner assessment.

AB-05 stops external release. The accessibility owner and release owner receive the preserved escape, including the automated result and manual review. Mitigation repairs the relation, then tests it with structured and human review. AB-06 stops the governance claim until status and applicability are corrected and qualified review occurs.

These are not six public disclosures. Internal notification is proportionate to the fixture. If a real release, person, platform, or regulated duty existed, the disclosure path would require new facts and qualified review. The course does not fabricate an external incident to make the case feel realistic.

When allow is reconsidered, it applies to a versioned release, not to the system forever. Record the owner’s accepted residual risk, monitoring triggers, evidence expiry, and conditions for reopening. An allow decision can be revoked by drift, new threat evidence, failed monitoring, or a changed component.

## Chapter 7 — Evidence preservation and response

The response sequence is detect, preserve, triage, contain, correct, notify, recover, and learn. Preservation occurs before cleanup changes the evidence.

Record the exact artifact or response, hashes, identifiers, timestamps, locale, account and tool state, source versions, model and parser versions if exposed, control output, reviewer observation, and known gaps. Preserve the command and environment for a deterministic lab. Use the minimum authorized personal or sensitive content; restrict access and document chain-of-custody limits.

A hash establishes file identity. It does not establish truth, authorship, legality, accessibility, or complete state. A screenshot may omit interface context. A log may omit a failed event. Keep missing evidence visible.

Containment limits further harm. Quarantine an unresolved record. Withhold an aggregate derived from dependencies. Pause release. Restore an accessible artifact. Correct a standards label. Do not destroy the original escape from the evidence package.

Correction must propagate. Updating one page is insufficient if the claim graph, structured representation, cached output, course material, benchmark label, or erratum remains stale. Name every dependent artifact and verify it.

For AB-05, preserve both the automated structure PASS and the manual observation that the claim-source association was inaccessible. If only the automated result survives, the record creates a false assurance. Notify the accessibility owner and release owner, repair the relation, perform manual and programmatic review, and keep external release stopped until recovery is verified.

[Pause. What is the immediate danger of correcting AB-05 before preserving the failing artifact? The team may be unable to explain the escape, validate the detector gap, or show that the repair addressed the observed failure.]

Negative and escaped results are evidence. Research integrity requires retaining them, including disagreements, deviations, failed checks, and version history.

## Chapter 8 — Notification, rollback, and recovery

Notification asks who needs to know, why, by when, under which policy or applicable requirement, and what evidence may be shared. Possible roles include system owner, evidence owner, privacy owner, accessibility owner, safety owner, governance owner, release owner, and affected party. External duties require qualified legal and sector review; this course is not legal advice.

Rollback is a restoration claim. Record trigger, authorized owner, exact target, pre-state identity, restoration action, success criterion, verification evidence, and residual effects. The sentence “we can revert” supplies none of these.

Each L07 control has a rollback specification. Restore the prior registry and rebuild the graph. Retain both dependency audit views. Restore the clean corpus and parser configuration. Restore an authorized redacted snapshot. Restore the last accessible control artifact. Restore the verified governance crosswalk. These are fixture contracts, not proof of external rollback.

Recovery validates the protected asset. After a registry restore, verify source joins and graph edges. After accessibility repair, verify claim-source relations programmatically and with assistive-technology review. After a status correction, verify source identity, date, status, jurisdiction, applicability, and ceiling.

Rollback can leave residual effects: notifications, caches, copied data, altered ordering, or incomplete logs. Record them. Test rollback inside authorized synthetic scope before depending on it. Ensure restoration does not erase incident evidence.

Close an incident only when correction is verified, dependent artifacts are updated, owners acknowledge the outcome, required notifications are complete, the control gap has an owner, and residual risk is recorded. Define reopening conditions if new evidence arrives.

[Pause. A rollback command returns success, but no one verifies the accessible relation. Which state exists? A tool return. Which state is missing? Recovery verification of ASSET-04.]

Build a recovery proof as a before-and-after table. The before column contains the frozen failing artifact, hash, reviewer observation, dependency list, and monitoring state. The action column contains the authorized restoration target and owner. The after column contains the new artifact hash, checks against the protected asset, dependent-artifact status, monitoring restart, and reviewer sign-off. A deviation column records anything that did not match the playbook.

For AB-01, successful recovery is not merely restoring a registry file. The rebuilt claim graph must contain resolvable source edges, and the audit must retain the quarantined attempt. For AB-02, dependency adjudication must change the aggregation without deleting the original evidence-family view. For AB-03, source-originated data must remain unable to become control state. For AB-04, the restored snapshot must be authorized and free of unapproved identifiers.

For AB-05, recovery needs the strongest mixed evidence: programmatic structure, keyboard navigation, assistive-technology relation review, and reviewer confirmation that each claim is associated with its source. For AB-06, the crosswalk must preserve source identity, exact status category, jurisdiction, as-of date, applicability note, and claim ceiling. A correct word alone is insufficient if its evidence has expired.

Notification closure also needs evidence. Record recipient, time, message class, acknowledgement, requested action, and completion status. Do not place secrets or unnecessary personal information in the notification log. If a duty or recipient is uncertain, escalate to the qualified owner rather than guessing publicly.

## Chapter 9 — L07 reproduction and decision

The validator runs the L07 script against six frozen inputs in a fresh temporary directory. The six cases are source identity, dependency amplification, instruction/content separation, privacy exposure, accessibility/attribution loss, and governance-force collapse.

The outcomes are three blocked, two detected, and one escaped. The escaped case is AB-05, with high residual severity and likely residual likelihood. The expected standard output says: `L07 PASS: 6 cases, 1 escape(s), release BLOCKED`.

Audit status PASS means the script processed the declared inputs, joins, risk vocabulary, and decision rule successfully. Release decision BLOCKED means an escaped high or critical residual case, or residual score above the declared threshold, stops release. The states answer different questions.

The run creates a threat model, case results, control coverage, governance-status audit, release decision, governance memo, rollback playbook, and run manifest. The W14 package validates all six input hashes and all eight output hashes. A fresh temporary directory prevents stale outputs from being mistaken for a new run.

The decision ceiling is synthetic control tests and governance-status vocabulary audit only. PASS is not certification. BLOCKED is not a claim that every external system is unsafe. The fixture is not legal advice, compliance evidence, or a production prevalence estimate.

[Pause. Complete: “The audit passed because the frozen process reproduced; release is blocked because AB-05 escaped the supplied control with high residual risk.” Then add: “No live attack or external platform claim follows.”]

The safe triage is: continue only the sandboxed course reproduction; mitigate AB-05; stop external release; preserve the escape; notify internal accessibility and release owners; assess any broader disclosure separately; validate recovery; and retain residual risk.

## Chapter 10 — Integrated release sentence

A defensible W14 conclusion names the system version, protected asset, actor, capability, entry point, trust boundary, harm, authorization, proportionality rationale, control evidence, detector coverage and false-positive limits, triage disposition, evidence location, notification owner, recovery criterion, and residual risk.

Here is the bounded version: “Within authorized offline system `SYS-RISK-DEMO-001`, six inert scenario classes cross two declared trust boundaries; supplied controls block three, detect two, and allow AB-05 to escape with high/likely residual risk. Detector evidence remains family-, domain-, threshold-, and denominator-bound, so it cannot establish intent or automatic punishment. The deterministic audit passes, but external release stops; the team preserves the accessible-relation escape, notifies the accessibility and release owners, mitigates and retests, validates recovery, and records residual risk. No live attack, platform mechanism, safety guarantee, compliance, certification, or legal conclusion follows.”

That sentence resists three common errors. It does not turn a detector into a judge. It does not turn a passed script into release approval. It does not turn a synthetic threat model into permission for live action.

[Final pause. Write your own three-sentence version. Sentence one defines the authorized defensive scope. Sentence two names the strongest control failure and detector ceiling. Sentence three states triage, owner, recovery, residual risk, and the no-live-claim boundary.]

Remember the method: model before testing; authorize before capability; choose the least intrusive route; distinguish prevention from detection; report family coverage and false positives; preserve before cleanup; notify with scope; verify rollback and recovery; keep residual risk visible; and stop when the boundary cannot be defended.
