# W14 Slide Script — Manipulation, Poisoning, and Defense

## Production conventions

- **Format:** 16:9, pale stone background, ink-black text, blue system boundaries, amber uncertainty, and red stop states. Labels and patterns duplicate every color meaning.
- **Typography:** minimum 30 pt body and 44 pt title. Keep public IDs visible when evidence roles are discussed.
- **Visual authorship:** all visuals are original instructional specifications. Do not use paper or TeX figures, toolkit interfaces, product screenshots, or operational attack illustrations.
- **Safety grammar:** scenario classes appear as abstract numbered paths only. No prompt, payload, credential, endpoint, evasion sequence, load pattern, or deployment procedure appears.
- **Accessibility:** every visual has exact alt text and a table or long-description equivalent. Dashed arrows mean plausible paths; solid arrows mean recorded fixture transitions; octagons mean stop.

## Slide 01 — When does optimization become system abuse?

**On-screen text**

> W14 · Manipulation, poisoning, and defense  
> Essential question: When does optimization become system abuse?

Footer: “Technical possibility is not authorization.”

**Visual specification**

An original balance places “defensive evidence value” on one side and “authorization, harm, reversibility” on the other. A boundary line separates a synthetic fixture from a hatched external system. A stop octagon guards the external side.

**Speaker notes**

Open by rejecting an intent-only boundary. A benign label does not make a test safe, and an adversarial scenario class is not automatically abusive inside an authorized synthetic fixture. The judgment attaches to evidence, scope, actors, capabilities, harms, and controls.

**Teaching check**

Learners name one fact that intent cannot establish.

**Alt text**

“A balance compares defensive evidence value with authorization, harm, and reversibility, while a stop sign blocks the path from a synthetic fixture to an external system.”

## Slide 02 — Five inspectable outputs

**On-screen text**

Produce:

1. versioned threat model;
2. authorization/proportionality card;
3. control lifecycle matrix;
4. detector coverage card; and
5. triage and incident decision.

**Visual specification**

Five stacked evidence cards use unique shapes: system map, signed scope rectangle, circular lifecycle, matrix grid, and five-branch decision fan. Each card shows a file-like identifier and review-status field.

**Speaker notes**

The course evaluates artifacts, not vocabulary recital. Each output can be inspected independently and linked to a version. A fluent threat narrative without authorization or evidence locations is incomplete.

**Teaching check**

Which output would reveal that a detector was never tested outside known templates?

**Alt text**

“Five distinct artifact cards represent a threat model, authorization card, lifecycle matrix, detector card, and triage decision.”

## Slide 03 — The seven-part boundary test

**On-screen text**

Truth · provenance · disclosure · system integrity · competitive fairness · authorization · proportionality

Intent alone does not decide.

**Visual specification**

Seven vertical gates form an inspection corridor. A proposed intervention card must pass each labeled gate. The intent label floats outside the corridor and has no shortcut arrow. Failed gates point to a stop or mitigation box.

**Speaker notes**

Truth includes implied claims. Provenance includes source dependence. Disclosure includes material interests. System integrity separates content from control. Competitive fairness covers fabricated consensus and unsupported harm. Authorization and proportionality govern the test itself.

**Teaching check**

Give a factually true edit that could still fail one other gate.

**Alt text**

“A proposed intervention must pass seven labeled gates, while intent remains outside and cannot bypass them.”

## Slide 04 — Model the system before threats

**On-screen text**

System = scope + components + flows + assets + actors + capabilities + entry points + trust boundaries + exclusions + owners

**Visual specification**

An original data-flow diagram shows ingest, normalize, retrieve, annotate, and release. A metadata band above lists scope and exclusions; a responsibility band below lists owners. No attack route is drawn yet.

**Speaker notes**

Starting with named attacks encourages coverage theater. System modeling exposes privacy, accessibility, dependency, and status harms that a narrow exploit list may miss. Version the diagram because changed components or permissions change the threat model.

**Teaching check**

What becomes stale when a new release component is added?

**Alt text**

“Five pipeline components sit between scope/exclusion and owner bands, demonstrating that the system is defined before any threat route.”

## Slide 05 — Assets and affected parties

**On-screen text**

L07 assets:

- source and citation integrity;
- claim-evidence lineage;
- privacy and authorization;
- accessible evidence presentation; and
- governance-status accuracy.

**Visual specification**

Five shield-shaped asset cards pair with affected-party labels: reader, source creator, researcher, reviewer, data subject, operator, assistive-technology user, and decision maker. Multiple connecting lines show shared impact.

**Speaker notes**

An asset is something worth protecting, not merely a file. Each harm should name an affected party. This makes accessibility and governance accuracy substantive system assets rather than presentation details.

**Teaching check**

Who is affected when duplicated evidence is mistaken for independent consensus?

**Alt text**

“Five protected assets connect to readers, creators, researchers, reviewers, data subjects, operators, accessibility users, and decision makers.”

## Slide 06 — Actors are bounded by capabilities

**On-screen text**

ACT-01 authorized learner  
ACT-02 independent reviewer  
ACT-03 untrusted corpus contributor

Capability grammar: verb + object + scope + time + approval

**Visual specification**

Three actor cards connect only to permitted action tokens. Disallowed tokens—credential, network, live target, persistence—sit outside a red boundary. Lines stop rather than showing attack procedures.

**Speaker notes**

Role labels do not imply unlimited access. Write “submit one inert synthetic record” rather than “corpus access.” The external capabilities are excluded, not latent possibilities for experimentation.

**Teaching check**

Rewrite “reviewer access” as a scoped capability.

**Alt text**

“Three actors connect to narrow permitted actions, while credential, network, live-target, and persistence capabilities remain outside the boundary.”

## Slide 07 — Entry points and trust boundaries

**On-screen text**

TB-01: untrusted corpus → ingest  
TB-02: analysis outputs → release

Entry points: ingest · normalize · retrieve · annotate · release

**Visual specification**

Five pipeline boxes cross two thick blue boundary lines. Boundary labels name required control families. Dashed abstract scenario paths end at controls; none includes payload content.

**Speaker notes**

At each entry, ask what may enter, who introduces it, which check occurs before the next zone, and what evidence proves the check ran. Unknown coverage stays visible as a hatched gap.

**Teaching check**

Which L07 boundary contains the accessible claim-to-source release problem?

**Alt text**

“The ingest and release transitions cross two labeled trust boundaries, with abstract scenario paths terminating at defensive controls.”

## Slide 08 — Threat paths terminate in harm

**On-screen text**

Scenario class → protected asset → affected party → harm → control → residual risk

Never end the diagram at “attack success.”

**Visual specification**

A six-column flow shows one inert example: dependency amplification reaches claim lineage, burdens a reviewer, distorts evidence weight, meets a dependency control, and retains medium residual risk. No mechanism for creating duplicates is shown.

**Speaker notes**

The diagram is defensive because it makes harm and control central. Scenario classes are enough. An operational sequence would exceed the course purpose and authorization.

**Teaching check**

Add an affected party and harm to the phrase “instruction/content boundary failure.”

**Alt text**

“An abstract scenario path proceeds through asset, affected party, harm, control, and residual risk rather than ending at attacker success.”

## Slide 09 — Authorization is a frozen record

**On-screen text**

Owner · system/version · actors · capabilities · synthetic inputs · dates · methods · retention · stop rules · rollback · exclusions

L07: course-owned files; no external traffic.

**Visual specification**

A signed authorization card is surrounded by hash, calendar, capability, storage, stop, and rollback fields. A scope bracket encloses only local synthetic components. The external region is visibly disconnected.

**Speaker notes**

Authorization expires when scope changes. Prior approval does not cover a new target, model, network, data class, or payload. A tool’s feature list is not permission.

**Teaching check**

Name one change that requires authorization review even if the research question stays the same.

**Alt text**

“A versioned authorization card binds owners, inputs, capabilities, retention, stops, and rollback to a disconnected synthetic scope.”

## Slide 10 — Proportionality chooses the least intrusive route

**On-screen text**

Evidence value versus:

harm · affected-party burden · reversibility · uncertainty · safer alternatives

Question: Can a static inert case answer it?

**Visual specification**

Three test-route cards—static review, disconnected simulation, external action—sit on an ascending risk stair. A selector chooses the first route that can answer the defensive question. The external route remains blocked.

**Speaker notes**

Interesting data do not justify disproportionate risk. Many control, schema, release-rule, and reviewer questions can be answered with frozen records. If no bounded route works, retain an evidence gap.

**Teaching check**

What evidence question in L07 requires no model execution at all?

**Alt text**

“A least-intrusive selector chooses static or disconnected tests before an external route, which remains blocked.”

## Slide 11 — Prohibited-action boundary

**On-screen text**

No live poisoning · evasion · credential use · unauthorized load · harmful payload · cloaking · deceptive attribution · third-party testing

Use inert scenario classes only.

**Visual specification**

Eight prohibited labels appear outside a large stop perimeter. Inside are only six numbered scenario cards and defensive questions. No arrows cross inward or outward.

**Speaker notes**

This is a hard course boundary. Learners do not need deployable mechanics to identify assets, controls, evidence, or response. Stop and move to L07 if a proposed activity approaches the perimeter.

**Teaching check**

Convert a request for a live demonstration into a safe synthetic question.

**Alt text**

“Eight prohibited action classes remain outside a stop perimeter containing only inert numbered scenarios and defensive questions.”

## Slide 12 — Prevention before the harmful transition

**On-screen text**

Examples:

stable source IDs · dependency register · content/config separation · synthetic-only allowlist · accessible table alternative · controlled status vocabulary

**Visual specification**

Six gate icons stand before the relevant pipeline transitions. Each gate has an evidence receipt below it. A note says “policy text is not implementation evidence.”

**Speaker notes**

Prevention acts before harm. Evidence must show the gate ran on the current version and reached its threshold. A control name or policy statement is not proof.

**Teaching check**

Which L07 prevention protects against a source-originated control action?

**Alt text**

“Six preventive gates appear before transitions, each requiring a separate implementation-and-test receipt.”

## Slide 13 — Detection after exposure is not prevention

**On-screen text**

Detection record:

family · domain · observation point · threshold · coverage · latency · FPR · review path

**Visual specification**

A detector sits after an exposure marker. A timing arrow labels prevention earlier and detection later. A coverage grid shows one covered and two unknown families. A denominator card sits beside FPR.

**Speaker notes**

Preserve temporal truth. L07 cases labeled detected were not blocked at the first gate. A detector needs a stated target family, denominators, and review process; silence about unknown families is not coverage.

**Teaching check**

Why can “we detected it” still describe a control failure?

**Alt text**

“A detector appears after exposure, with separate coverage, false-positive denominator, and human-review records.”

## Slide 14 — PAPER-37: strong known-family results, weak external coverage

**On-screen text**

PAPER-37 · SCI-Defense

Read together:

- some high known-template detection results;
- differing reported false-positive cells; and
- very weak results for several new black-box families.

**Visual specification**

An original family-by-detector matrix uses printed values only as qualitative bands: high, mixed, near-zero. A bracket separates known templates from held-out families. A large average score is crossed out.

**Speaker notes**

Do not reproduce attack mechanics. The scientific lesson is that known-template performance cannot stand in for strategy-external defense. Retain sample sizes, family labels, and the difference between aggregate and independent false-positive estimates.

**Teaching check**

Which result would you show first when evaluating claimed generalization?

**Alt text**

“A detector matrix separates strong known-template cells from weak held-out-family cells and rejects one aggregate score.”

## Slide 15 — PAPER-40: aggregation can fail under constructed contamination

**On-screen text**

PAPER-40 · controlled evidence-aggregation study

Use for: per-condition robustness and dependency questions  
Ceiling: no universal winner; repetition, implementation, and uncertainty gaps remain

**Visual specification**

Three evidence-family cards flow into several generic aggregators, then a claim decision. Dependency links join two cards. A result table contains “varies by condition” rather than operational inputs.

**Speaker notes**

The source motivates family-aware failure matrices, source-dependency analysis, abstention, and repeated runs. W14 does not reproduce poisoning procedures or declare an aggregator universally robust.

**Teaching check**

Why can five pages derived from one source not count as five independent confirmations?

**Alt text**

“Correlated evidence families feed generic aggregators whose outcomes vary by condition, with no universal winner.”

## Slide 16 — Extension routes form a defense matrix

**On-screen text**

PAPER-07 controlled adversarial search  
PAPER-11 local dynamics  
PAPER-13 citation vulnerabilities  
PAPER-34 bounded persuasion mechanisms  
PAPER-39 recommendation-agent risks  
PAPER-41 detection benchmark

**Visual specification**

Six source-role cards occupy columns labeled behavior, dynamics, citation, mechanism, recommendation, and detection. A horizontal ceiling bar says “setting-specific evidence; no live authorization.”

**Speaker notes**

These sources extend the question; they do not merge into one leaderboard. PAPER-11 is theoretical, PAPER-34 is task/model bounded, and PAPER-39 is benchmark bounded. PAPER-07 and PAPER-13 do not establish universal production vulnerability.

**Teaching check**

Choose one route and state its role plus one forbidden extrapolation.

**Alt text**

“Six extension papers occupy distinct threat-evidence columns beneath a common setting-specific ceiling.”

## Slide 17 — PAPER-41: detection is not accusation

**On-screen text**

Detected style ≠ false claim ≠ malicious intent ≠ policy violation ≠ legal wrongdoing

Route to review; preserve appeal and correction.

**Visual specification**

A detector score enters a private review queue. Four dashed arrows toward accusation, suppression, punishment, and public label are blocked. A correction and appeal loop returns to the review queue.

**Speaker notes**

PAPER-41 supports dataset-bound detection evaluation. Ground truth and coverage limits prevent use as automatic proof. Base rate and false-positive burden matter, especially when the downstream action is high impact.

**Teaching check**

What additional evidence is needed before a public claim about intent?

**Alt text**

“A detector score routes to private review while automatic accusation, suppression, punishment, and public labeling are blocked.”

## Slide 18 — Controls form a lifecycle

**On-screen text**

Prevent → detect → triage → contain → preserve → correct → notify → recover → learn

Residual risk remains.

**Visual specification**

A circular lifecycle has nine labeled segments. A residual-risk register sits beside rather than inside the loop. Stop gates appear at authorization, privacy, accessibility, logging, and rollback failures.

**Speaker notes**

A detector is one segment. Response needs named owners and evidence. Learning does not close the risk register. Each transition should leave a trace and have a time expectation.

**Teaching check**

Which event must occur before cleanup can destroy evidence?

**Alt text**

“A nine-stage control lifecycle is accompanied by a separate residual-risk register and five stop gates.”

## Slide 19 — Preserve evidence before changing state

**On-screen text**

Preserve:

artifact · hash · time · locale/account/tool state · source/model/parser versions · control output · reviewer observation · known gaps

**Visual specification**

An evidence bundle contains immutable labeled cards and a chain-of-custody ledger. A privacy minimization filter sits at acquisition. A warning says “hash proves identity, not truth.”

**Speaker notes**

Evidence preservation and privacy must coexist. Capture the smallest authorized evidence needed, restrict access, and document gaps. A screenshot or log may omit relevant state.

**Teaching check**

Name one field needed to reproduce a time-sensitive observation and one field a hash cannot establish.

**Alt text**

“A minimized evidence bundle records identity, context, versions, control output, and gaps while warning that hashes do not prove truth.”

## Slide 20 — Notification and recovery have owners

**On-screen text**

Notify whom? owner · privacy · accessibility · evidence · release · affected party

Recover what? protected asset + dependent artifacts + monitoring

**Visual specification**

A triage hub routes to six role cards through decision gates. A recovery checklist spans source, graph, accessible relation, status label, dependent materials, and monitoring. Public disclosure is a separate qualified-review branch.

**Speaker notes**

Notification depends on harm, contract, policy, jurisdiction, and qualified review. It is not automatically public. Recovery propagates corrections and confirms the protected asset, not only a tool return.

**Teaching check**

Who owns AB-05 notification and which artifact must be verified after repair?

**Alt text**

“A triage hub routes scoped notifications to named owners and a recovery checklist verifies both the protected asset and dependent artifacts.”

## Slide 21 — Rollback is a tested restoration claim

**On-screen text**

Trigger · owner · exact target · pre-state · action · success criterion · verification · residual effects

“We can revert” is not evidence.

**Visual specification**

A before-state hash and a recovery target bracket a rollback operation. A verification gate checks six asset properties. Residual effects flow into a register rather than disappearing.

**Speaker notes**

Rollback can leave notifications, caches, copied data, or lost ordering. Test it in the authorized synthetic environment. Preservation must survive restoration.

**Teaching check**

What proves that restoring a registry also repaired the claim graph?

**Alt text**

“A rollback moves from a frozen pre-state toward a recovery target, then passes verification while residual effects remain recorded.”

## Slide 22 — Five safe triage dispositions

**On-screen text**

ALLOW · SANDBOX · MITIGATE · STOP · DISCLOSE

Disposition = evidence + authorization + harm + reversibility + owner

**Visual specification**

Five branches leave one review diamond. Allow has a monitoring loop; sandbox stays inside isolation; mitigate loops through retest; stop ends at an octagon; disclose reaches a scoped notification gate. No branch equates detection with accusation.

**Speaker notes**

Dispositions can form a sequence. Sandbox can expose a gap, mitigation can fail, release can stop, and an owner can be notified. Disclose is an assessment under policy and qualified review, not a reflexive public post.

**Teaching check**

Triage a high-score, low-ground-truth detector alert without assuming guilt.

**Alt text**

“Five labeled triage branches show monitoring, isolated testing, mitigation and retest, stopping, and scoped disclosure review.”

## Slide 23 — L07: six cases, one escape, release blocked

**On-screen text**

Outcomes: blocked 3 · detected 2 · escaped 1  
AB-05: accessible claim-source relation lost  
Audit status: PASS  
Release decision: **BLOCKED**

**Visual specification**

Six scenario cards enter two independent result lanes. The process lane reaches PASS after deterministic checks. The release lane stops at AB-05. A table alternative lists all six case IDs and results.

**Speaker notes**

PASS describes correct fixture processing. BLOCKED applies the release rule to escaped high residual risk. The result is not certification, compliance, legal advice, or a safety guarantee.

**Teaching check**

Explain the two statuses in one sentence without calling them contradictory.

**Alt text**

“Six inert cases pass deterministic audit processing, but the release lane is blocked by escaped accessibility case AB-05.”

## Slide 24 — The W14 release sentence

**On-screen text**

Name system · asset · actor · capability · boundary · harm · authorization · proportionality · control · detector limit · owner · recovery · residual risk

Sources: PAPER-37 · PAPER-40 · PAPER-07 · PAPER-11 · PAPER-13 · PAPER-34 · PAPER-39 · PAPER-41 · G08 · S01 · S02

**Visual specification**

Thirteen labeled evidence blocks support one narrow release sentence. A source strip distinguishes core, extend, tool candidate, and standards/guidance candidates by borders and text. Three stop octagons mark live action, missing authorization, and failed rollback.

**Speaker notes**

Close with the fixture decision: stop external release, preserve AB-05 evidence, notify named owners, repair and retest the accessible relation, validate recovery, and retain residual risk. No live attack occurred and no production system is characterized.

**Teaching check**

Exit ticket: write one bounded release sentence containing a detector ceiling and triage disposition.

**Alt text**

“Thirteen evidence blocks support a bounded release sentence, with public source roles and stop signs for live action, missing authorization, and failed rollback.”
