# W11 Slide Script — Multilingual, Multimodal, and Agentic Surfaces

## Production conventions

- **Format:** 16:9, warm white field, midnight blue structure, teal evidence marks, amber uncertainty, and red stop states. Color is always paired with labels, shapes, or line styles.
- **Typography:** at least 30 pt body and 44 pt title. Keep public identifiers such as PAPER-02 and S04 visible when a source role is discussed.
- **Visual authorship:** every diagram below is an original course visual specification. Do not paste paper figures, TeX figures, product screenshots, or standards diagrams. Use only course-owned synthetic or separately licensed production assets.
- **Boundary grammar:** solid arrows are recorded transformations; dashed arrows are hypotheses; double outlines are authorized states; octagons are stops; hatched regions are unknown.
- **Accessible production:** retain the exact alt text, add table alternatives for data-bearing visuals, and never require animation, position, or color alone to understand a relation.

## Slide 01 — When the source changes form or the system can act

**On-screen text**

> W11 · Multilingual, multimodal, and agentic surfaces  
> Essential question: What changes when the source is an image or the system can act?

Footer: “Representation is not identity. Recommendation is not action.”

**Visual specification**

Split the canvas. Left: one synthetic chart fans into seven labeled representation cards. Right: a speech bubble ends before a locked tool gate. A central vertical rule labels the two discontinuities, representation and side effect. No interface chrome appears.

**Speaker notes**

Open with two questions. If a system repeats a chart caption, did it inspect the pixels? If it says “I recommend creating a draft,” did anything change outside the conversation? Both questions expose a hidden collapse. This week supplies ledgers that keep transformations and actions observable.

**Teaching check**

Learners write one piece of evidence needed to answer each opening question.

**Alt text**

“A chart branches into seven representations, while a recommendation stops at a locked action gate, illustrating separate representation and action problems.”

## Slide 02 — Outcomes leave inspectable artifacts

**On-screen text**

By the end, produce:

1. a representation ledger;
2. evidence-bearing alt and a long description;
3. a locale/model/pooling record;
4. an action-state audit; and
5. a bounded release decision.

**Visual specification**

Five numbered artifact cards form a staircase. Each card carries a text label plus an original geometric glyph: linked squares, a text block, a globe grid, a permission key, and a stop octagon. Numbers establish sequence.

**Speaker notes**

Explain that every outcome is inspectable. Learners are not rewarded for saying a system “understands images” or “is agentic.” They must leave a ledger, description, comparison record, action trace, and conclusion whose ceiling follows from observable evidence.

**Teaching check**

Ask which artifact would reveal that two language observations used different model versions.

**Alt text**

“Five ordered cards represent the representation ledger, accessible description, locale record, action audit, and release decision.”

## Slide 03 — One asset, many representations

**On-screen text**

Asset bytes → pixels → regions  
Page context → alt · caption · metadata  
Pixels → OCR · embedding  
Retrieved material → generated description

Rule: “Name the representation you observed.”

**Visual specification**

An original node-link diagram uses rectangles for stored artifacts and hexagons for transformations. Solid edges are recorded; a dashed final edge points to a closed-surface answer with a hatched interior. Every node has a short identifier.

**Speaker notes**

Stress that an asset, its enclosing page, and derived representations are distinct audit units. A correct answer that matches alt text does not prove pixel processing. The diagram is an inspection grammar, not a reverse-engineered commercial architecture.

**Teaching check**

Give “the caption and answer share a phrase.” Learners state one supported observation and one unsupported mechanism claim.

**Alt text**

“Stored asset and page artifacts pass through separate pixel, OCR, embedding, retrieval, and description transformations before a hatched unknown answer surface.”

## Slide 04 — The representation ledger

**On-screen text**

| Field | Record |
|---|---|
| Input | ID + hash |
| Transform | name + version |
| Context | crop + locale + parameters |
| Output | ID + hash |
| Review | supports / incomplete / contradicts / unknown |

**Visual specification**

A five-row ledger fills the slide. One example path is shown in a side rail: synthetic asset → OCR vX → token output. A red outline around the missing version demonstrates a failed record. Provide the table as real text.

**Speaker notes**

Hashes establish byte identity, not semantic correctness. Transformer names require versions and configuration. If identity is not exposed, record unknown. Do not invent a familiar OCR engine, captioner, embedding model, or hidden retrieval design.

**Teaching check**

Which field helps reproduce an OCR failure, and which field helps judge its semantic effect?

**Alt text**

“A ledger records input, transformer, context, output, and review, with a missing transformer version marked as a failure.”

## Slide 05 — What the synthetic chart actually supports

**On-screen text**

Fictional fixture:

- Group A: 12 verified cases
- Group B: 8 verified cases
- Unresolved: 4 records
- No intervention or causal claim

**Visual specification**

Draw an original two-bar chart with labeled values 12 and 8. To the right, a dotted-outline box reads “4 unresolved.” Axes say “fictional group” and “verified cases.” A footer labels the image course-owned synthetic data.

**Speaker notes**

This is the only quantitative visual used in the worked case. State the claim ceiling before transformations begin. The display supports counts in a fictional fixture. It does not support “performance,” an intervention effect, a population estimate, or any real-world conclusion.

**Teaching check**

Learners reject the sentence “Group A performed 50 percent better because of the intervention” field by field.

**Alt text**

“Fictional bar chart with 12 verified cases for Group A, 8 for Group B, and a separate note that four records are unresolved; no causal meaning is shown.”

## Slide 06 — Short alt carries task-relevant evidence

**On-screen text**

Weak: “Chart.”  
Incomplete: “Bar chart comparing two groups.”  
Evidence-bearing: “Bar chart: Group A has 12 verified cases and Group B 8; four records remain unresolved.”

**Visual specification**

Three horizontal text panels align with the same small outline of the chart. The first has one empty evidence cell, the second two empty cells, and the third all value and uncertainty cells filled. Labels, not color, indicate weak, incomplete, and evidence-bearing.

**Speaker notes**

Short alt should let a nonvisual learner perform the teaching task. It need not repeat every surrounding sentence, but it must preserve the values and uncertainty that matter. Decorative null alt is inappropriate when the image carries core evidence.

**Teaching check**

Ask learners to identify the minimum additional phrase that turns the second version into task-equivalent evidence.

**Alt text**

“Three alt-text examples progress from only naming a chart to carrying both group values and the unresolved-record uncertainty.”

## Slide 07 — Long description reveals structure and limits

**On-screen text**

A long description states:

- reading order and encodings;
- exact labels, values, and units;
- annotations and uncertainty;
- observed pattern versus interpretation; and
- a table alternative when values matter.

**Visual specification**

Place the chart on the left and an ordered description skeleton on the right. Fine leader lines connect “bar height,” “value label,” and “unresolved note” to matching prose. The final line ends at an octagon reading “no causality.”

**Speaker notes**

Long descriptions are not licenses to narrate beyond the visual. They expose how a visual is read and where interpretation must stop. The table alternative should be available as actual structured content, not hidden inside an image.

**Teaching check**

Which sentence belongs in a long description but would overload short alt?

**Alt text**

“A chart is connected to an ordered long-description outline covering encoding, values, uncertainty, and a no-causality stop.”

## Slide 08 — OCR is a fallible transformation

**On-screen text**

Frozen OCR output: `Group A 12; Group B; unresolved 4`

Expected visual evidence: `12`, `8`, `4`

Disposition: **incomplete; preserve the failure**

**Visual specification**

Show three token boxes from the visual evidence and an OCR rail below. The 8 box maps to an empty dashed slot. An audit tag says “do not silently repair.” Reading order is numbered.

**Speaker notes**

OCR can drop a small label, reorder columns, or confuse scripts. Its output must remain distinguishable from authored alt and source data. Correcting the frozen output without a new version would erase evidence about the failed transformation.

**Teaching check**

Does the OCR contradict the value 8 or merely fail to carry it? Defend one label.

**Alt text**

“OCR carries 12 and 4 but has an empty slot where the visual value 8 should appear, so it is labeled incomplete and preserved.”

## Slide 09 — Caption agreement does not prove pixel use

**On-screen text**

PAPER-25 · caption-in-context proxy

Observed agreement may involve:

- caption text;
- surrounding page text;
- pixels;
- OCR; or
- an unobserved combination.

**Visual specification**

Five possible input paths converge on an answer card. Only caption agreement is solid; the mechanism paths are dashed and end behind a hatched system boundary. A source-role badge says “bounded proxy, not universal rule.”

**Speaker notes**

Use PAPER-25 to teach evidence ceilings. Its context can motivate a caption pathway, but it does not justify direct-image-consumption claims, a universal caption rule, or transfer to an untested system. Similar text is evidence of similarity, not internal causation.

**Teaching check**

Rewrite “the model saw the image” as a defensible closed-surface observation.

**Alt text**

“Caption, page text, pixels, OCR, and an unknown combination are possible paths to an answer, while only caption-answer agreement is observed.”

## Slide 10 — Embedding and retrieval are not explanations

**On-screen text**

Embedding record = encoder + version + input + preprocessing + vector policy  
Retrieval record = query + index snapshot + candidate IDs + ranks/scores

Neither is a citation or human-readable claim.

**Visual specification**

An image card enters a numeric-vector capsule, then an index grid, then three candidate cards. A parallel evidence lane requires accessible text and source identity. The lanes remain separate until a review gate.

**Speaker notes**

An embedding can support retrieval without preserving an interpretable explanation. Record its identity and downstream behavior. A correct candidate can still be described incorrectly or attributed to the wrong source. Do not use vector proximity as entailment.

**Teaching check**

Name one record needed to diagnose candidate loss and one needed to diagnose attribution error.

**Alt text**

“An image becomes a vector and retrieval candidates in one lane, while accessible evidence and source identity remain a separate review lane.”

## Slide 11 — Cross-modal claim matrix

**On-screen text**

| Representation | 12 | 8 | 4 unresolved | Causality |
|---|---|---|---|---|
| Pixels/table | supports | supports | supports | absent |
| English alt | supports | supports | supports | absent |
| Caption | incomplete | incomplete | absent | absent |
| OCR | supports | absent | supports | absent |
| Generated description | supports | supports | absent | contradicts ceiling |

**Visual specification**

Render the table as text with words and simple shapes in every status cell. Add a callout to “contradicts ceiling,” explaining that the word “because” exceeds evidence. Do not rely on traffic-light color.

**Speaker notes**

Compare identity, values, units, uncertainty, time, and causal language rather than compressing everything into one consistency score. A representation can be incomplete without contradiction. Generated prose can contain correct numbers and still add an unsupported mechanism.

**Teaching check**

Which row could answer a count question but not an uncertainty question?

**Alt text**

“A matrix shows which representations carry 12, 8, unresolved 4, and causality; OCR misses 8 and generated text exceeds the causal ceiling.”

## Slide 12 — PAPER-02 defines a threat class, not a universal weakness

**On-screen text**

PAPER-02 · coordinated image/text manipulation in evaluated VLM ranking conditions

Use for: defensive threat questions  
Do not infer: every VLM is vulnerable · a commercial mechanism · permission to attack

**Visual specification**

An original threat card sits inside a bounded test frame. Three arrows attempting to leave the frame hit labeled stops: universality, hidden mechanism, live reproduction. A defensive checklist points inward to discrepancy logs and release gates.

**Speaker notes**

The source motivates checking whether coordinated representations can evade or dominate controls in the evaluated class. W11 does not reproduce attacks. It translates the threat into defensive auditing: compare representations, preserve provenance, detect disagreement, and stop when evidence cannot be verified.

**Teaching check**

Turn one offensive-sounding question into a safe audit question about the synthetic fixture.

**Alt text**

“A bounded PAPER-02 threat frame blocks universal, hidden-mechanism, and live-attack claims while allowing defensive discrepancy checks.”

## Slide 13 — S04 integrity has a strict ceiling

**On-screen text**

S04 can support a bounded provenance-integrity statement.

Integrity ≠ truth ≠ authorship ≠ rights ≠ accessibility ≠ ranking

**Visual specification**

Six equal columns form a judgment rack. Only “integrity” connects to a manifest-and-hash symbol. Five other columns remain separate, each with a question mark and its own required review badge. Equality signs are visibly crossed out between columns.

**Speaker notes**

S04 is the controlled standards candidate for the C2PA technical specification. Tamper-evident binding and manifest validation do not decide whether content is true, who authored it, whether use is licensed, whether evidence is accessible, or whether it affected ranking.

**Teaching check**

Give an intact but false synthetic image. Which judgment passes, and which fails?

**Alt text**

“Six separate judgment columns show that provenance integrity connects to a manifest, while truth, authorship, rights, accessibility, and ranking require other evidence.”

## Slide 14 — Multilingual work begins with entity identity

**On-screen text**

Entity ledger:

canonical ID · native name · aliases · transliteration · translated label · homonyms · geography · valid time · reviewer

**Visual specification**

A central fictional entity ID connects to three script-specific name cards, two alias cards, and one homonym warning. Transliteration edges are dotted; verified identity edges are solid. A reviewer stamp is pending on the translated label.

**Speaker notes**

Translation and transliteration solve different problems. Neither guarantees identity. A culturally specific label may shift construct or connotation. Use fictional names in the exercise, and preserve original script plus a stable canonical identifier.

**Teaching check**

Why is a string match insufficient when two languages use different scripts or share a homonym?

**Alt text**

“A canonical fictional entity links to native-script names, aliases, transliterations, and a homonym warning, with translation review pending.”

## Slide 15 — Locale is part of the observation cell

**On-screen text**

Record: surface · observable version · query · language · locale · geography · account · time · repetition · search/tool state · transformer/model identity

Unknown is a value, not a blank.

**Visual specification**

Two comparison cells, `en-US` and `zh-CN`, show a grid of the same fields. Three fields differ and are outlined individually. A bracket above the cells says “more than language changed.”

**Speaker notes**

Language comparisons are confounded if locale, geography, account state, time, model, or tool state also changes. Record these fields at observation time. Do not retroactively infer a model identity from marketing names or assume a stable service version.

**Teaching check**

Learners name two simultaneous changes that would block a language-only interpretation.

**Alt text**

“English-US and Chinese-China observation cells differ in several recorded fields, showing that language alone cannot explain the contrast.”

## Slide 16 — Pooling is a scientific decision

**On-screen text**

Pool only after checking:

construct · task equivalence · entity alignment · sampling frame · model state · scoring rule · missingness · heterogeneity

Fixture decision: **pooling blocked pending bilingual and measurement-equivalence review**

**Visual specification**

Two language streams approach a pooling basin. Eight labeled gates stand before the basin; two are open and six are pending. A stop octagon prevents a pooled mean from appearing.

**Speaker notes**

A pooled number can be arithmetically correct and scientifically meaningless. The fictional fixture records both language rows but does not authorize aggregation. Back-translation alone is not cultural validity; affected-language review and construct equivalence are required.

**Teaching check**

Which missing check would most threaten interpretation in a culturally specific recommendation task?

**Alt text**

“Two language streams cannot enter a pooling basin because most equivalence gates remain pending.”

## Slide 17 — PAPER-01 is a quarantine lesson

**On-screen text**

PAPER-01 · time provenance unresolved relative to the 2026-08-24 freeze

Allowed: audit the inconsistency  
Blocked: import substantive findings as course facts

**Visual specification**

A document card sits inside a transparent quarantine box. A timeline shows course freeze before reported collection markers. One magnifier arrow enters for metadata review; one evidence arrow is stopped from entering the lesson claims.

**Speaker notes**

Quarantine does not assert that every reported finding is false. It states that the current evidence package cannot pass admission because its time record is unresolved. The proper response is preserve, investigate, and exclude substantive use until resolved.

**Teaching check**

Write a sentence that distinguishes quarantine from retraction or falsification.

**Alt text**

“PAPER-01 is enclosed in a quarantine box because timeline markers conflict with the course freeze; metadata audit is allowed but claim admission is blocked.”

## Slide 18 — PAPER-05 is an architecture audit, not causal proof

**On-screen text**

PAPER-05 · industrial VLM/agent architecture case

Audit: units · denominators · randomization · intervals · attribution · missing uncertainty

Do not claim: reproduced traffic or user growth · causal business effect

**Visual specification**

An architecture sketch is represented only by original generic boxes, overlaid with six audit lenses. A separate business-outcome arrow ends at a locked evidence gate. No numerical marketing result appears.

**Speaker notes**

Use the paper to practice reading industrial architectures and evidence gaps. Reported outcomes remain tied to their source and are not reproduced in W11. Architecture plausibility is not causal identification, and a commercial case does not reveal every agent’s hidden pipeline.

**Teaching check**

What additional design evidence would be required before a business-growth causal claim?

**Alt text**

“A generic industrial architecture receives six audit lenses, while a business-effect arrow is blocked at an evidence gate.”

## Slide 19 — A recommendation stops before the action boundary

**On-screen text**

“Draft a correction memo” = recommendation

It does not establish:

tool proposal · permission · attempt · return · verified outcome · rollback

**Visual specification**

A text bubble sits left of a thick boundary. On the right are six unfilled action-state boxes. There is no connecting arrow. A note says “absence of transition is meaningful.”

**Speaker notes**

Agentic evaluation begins by refusing to treat action language as action evidence. A structured proposal is still not execution. A tool return is still not an independently verified outcome. Each transition needs its own logged event.

**Teaching check**

Ask for the minimum event that would distinguish a recommendation from an attempted operation.

**Alt text**

“A recommendation bubble remains separated from unfilled proposal, permission, attempt, return, verification, and rollback states.”

## Slide 20 — The seven-state action ledger

**On-screen text**

Recommend → propose → authorize → attempt → return → verify → roll back

At each edge: actor · capability · target · time · evidence

Rule: “No state implies the next.”

**Visual specification**

Seven boxes form a horizontal state machine. Every arrow passes through a small evidence diamond. The authorization box has a double outline. The return-to-verify arrow is dashed until independent evidence is attached.

**Speaker notes**

Authorization binds exact scope and arguments. Attempt records invocation, not success. Return records what a tool said, not necessarily what occurred. Verification observes the intended and unintended outcome. Rollback is another action and must itself be verified.

**Teaching check**

Why is a success response insufficient to mark the verification box complete?

**Alt text**

“Seven action states are separated by evidence diamonds; authorization is specially marked, and tool return does not automatically reach verification.”

## Slide 21 — Permissions and side effects before execution

**On-screen text**

Before attempt:

- least-privilege capability;
- exact target and argument binding;
- side-effect preview;
- idempotency or deduplication;
- durable logging;
- stop and rollback conditions.

**Visual specification**

A permission key fits only one synthetic target lock. Around it, six labeled guardrail cards form a ring. External accounts and networks sit outside a disconnected boundary with crossed connectors.

**Speaker notes**

Separate read from write, draft from publish, and synthetic buffer from external account. Side effects include notifications, costs, privacy exposure, downstream triggers, and duplicate writes. Permission is not transferable from a broad conversation-level intent.

**Teaching check**

Which permission distinction prevents a draft-only exercise from becoming a publication action?

**Alt text**

“A least-privilege key opens only a synthetic target, surrounded by guards for scope, effects, idempotency, logs, stops, and rollback.”

## Slide 22 — PAPER-39 bounds recommendation-agent risk claims

**On-screen text**

PAPER-39 · recommendation-agent risk and defense benchmark

Rates belong to: evaluated suite + agents + attack variants + conditions

Not: universal vulnerability · guaranteed defense · permission for live testing

**Visual specification**

A benchmark frame contains four labeled condition cards. Report arrows terminate at the frame. Outside, three oversized universal claims are crossed out. A safe arrow leads to permission and containment questions.

**Speaker notes**

Use PAPER-39 to motivate threat modeling, scoped permissions, logging, and verification. Preserve the benchmark’s evaluated context. W11 does not run attacks, connect tools, or claim that a commercial recommendation agent shares the same internal design.

**Teaching check**

Complete: “PAPER-39 supports a question about ___, but not a claim that ___.”

**Alt text**

“PAPER-39 findings remain inside a bounded benchmark frame, while universal vulnerability, guaranteed defense, and live-testing claims are rejected.”

## Slide 23 — L07: audit PASS, release BLOCKED

**On-screen text**

L07 deterministic result:

- 6 abuse cases
- blocked 3 · detected 2 · escaped 1
- blocker: AB-05 accessible claim-to-source association
- audit status PASS
- release decision **BLOCKED**

**Visual specification**

A two-lane decision graphic separates process status from release status. The upper lane reaches PASS after fixture checks. The lower lane encounters one escaped high-residual AB-05 card and reaches a red labeled stop. Counts are printed in both lanes.

**Speaker notes**

These states are compatible. PASS means the deterministic audit processed the declared fixture and rules. BLOCKED means the escape meets the release stop rule. Neither is certification, a compliance determination, legal advice, or a general safety verdict.

**Teaching check**

Learners explain the two lanes without using the word “contradiction.”

**Alt text**

“A process lane passes deterministic checks, while a separate release lane is blocked by escaped accessibility case AB-05.”

## Slide 24 — The W11 release sentence

**On-screen text**

Name the representation.  
Name the locale and model state.  
Name the action state.  
Separate integrity, truth, rights, accessibility, and ranking.  
Stop on unknown authorization, unverifiable evidence, or failed rollback.

Sources: PAPER-02 · PAPER-25 · PAPER-39 · PAPER-05 audit-only · PAPER-01 quarantined · S04 bounded

**Visual specification**

Five vertical evidence pillars support a narrow conclusion beam. Three stop octagons guard the exit. A source-role strip uses labels and border patterns to distinguish core, audit-only, quarantined, and standards-candidate routes.

**Speaker notes**

Close with the bounded case conclusion. The fictional asset shows how transformations can omit or overstate evidence. The bilingual comparison remains unpooled. The action trace contains no live action. L07 blocks release because accessibility evidence failed. Invite learners to prefer an explicit unknown or stop over a fluent story about hidden machinery.

**Teaching check**

Exit ticket: one multimodal claim, one multilingual claim, and one action claim, each with observable evidence and a stop condition.

**Alt text**

“Five evidence pillars support a bounded conclusion, while stop signs prevent release when authorization, evidence, or rollback is unresolved; source roles are explicitly separated.”
