# W11 Worked Case — A Bilingual Chart and a Disconnected Action Trace

## Case status

This case is a frozen, course-owned, synthetic exercise. It contains no person, real organization, external account, live endpoint, production agent, or platform observation. Identifiers beginning `synthetic://` are inert labels. The chart is specified in text and is not copied from a paper, TeX source, website, or product screenshot. All proposed diagrams are original course designs.

The case has two connected questions. First, which representation carries which evidence from a fictional bilingual chart? Second, what additional records are required when a system moves from recommending text toward a tool action? The release decision also uses deterministic Lab L07. A successful local audit does not establish general safety, accessibility conformance, compliance, certification, legal adequacy, or a live-system effect.

Controlled sources keep distinct roles. PAPER-02 and PAPER-25 inform bounded multimodal threat and caption-proxy questions. PAPER-39 informs bounded recommendation-agent risk questions. PAPER-05 is audit-only and supplies no causal business effect. PAPER-01 is quarantined for unresolved time provenance and supplies no substantive finding. S04 has an integrity-only ceiling; it does not establish truth, authorship, rights, accessibility, or ranking.

## 1. Frozen question, units, and stop conditions

**Question:** Can the fictional evidence package support a count comparison in English and Chinese, and can the action trace support a claim that an external action occurred?

The asset unit is `ASSET-W11-001`, a fictional two-bar chart. The claim unit is one proposition about a count, unresolved status, or cause. The representation unit is a frozen string or structured record produced at one transformation stage. The locale unit is a fully specified observation cell. The action unit is one logged state transition, not an entire conversation.

Stop immediately if an asset hash changes, a rights field is missing, a real or personal asset is introduced, a core claim is unavailable in an accessible representation, a language row lacks a locale/model record, cross-locale pooling is performed before review, an action lacks exact authorization, the target is external, a side effect is not logged, or rollback is unavailable. Stop also if L07 input hashes drift or the reproduced decision differs from the frozen contract.

The estimand language is deliberately modest. We are not estimating a causal effect. We are auditing whether declared representations support specific fictional propositions and whether action evidence reaches a named state.

## 2. Asset and source record

The asset declaration is:

| Field | Frozen value |
|---|---|
| Asset ID | `ASSET-W11-001` |
| Origin | course-authored synthetic fixture |
| Media design | two labeled bars plus unresolved-record note |
| License | course-owned synthetic, classroom reuse permitted |
| Personal data | none |
| Group A | 12 verified cases |
| Group B | 8 verified cases |
| Unresolved | 4 records |
| Causal status | no intervention, temporal, or causal information |

The imagined visual uses midnight blue for both bar outlines, different hatch patterns, printed value labels, and an explicit table alternative. Color is redundant. The long description supplies reading order and claim limits. This description is the production specification; no bitmap is required to reproduce the reasoning exercise.

The evidence supports “A has four more verified cases than B” and, with B as the explicit denominator, the arithmetic statement “A’s count is 50 percent higher than B’s count.” It does not support “A performed better,” because “performance” is undefined. It does not support “the intervention caused improvement,” because there is no intervention or design. Four unresolved records make the completeness boundary visible; they must not be assigned to either group.

## 3. Representation ledger

The manifest freezes these representation roles:

1. `REP-PIXEL-EVIDENCE` is a text declaration of what the course-owned visual encodes: A 12, B 8, unresolved 4, and no causality.
2. `REP-ALT-EN` is evidence-bearing English alt: “Fictional bar chart: Group A has 12 verified cases and Group B has 8; four records remain unresolved.”
3. `REP-ALT-ZH` is a Chinese course translation: “虚构柱状图：A组有12个已核实案例，B组有8个；另有4条记录尚未解决。” It is marked pending affected-language review.
4. `REP-CAPTION-EN` says only, “Group A has more verified cases than Group B in the fictional sample.” It preserves direction but omits magnitudes and unresolved records.
5. `REP-OCR-EN` says, “Group A 12; Group B; unresolved 4.” It deliberately omits 8.
6. `REP-LONG-EN` describes both bars, exact values, hatch encodings, the unresolved note, and the no-causality ceiling.
7. `REP-EMBED-CONFIG` records only an offline placeholder encoder identity and parameters; no vector is treated as human-readable evidence.
8. `REP-RETRIEVAL` records that the synthetic asset is the first and only candidate for the fixture query. It does not establish entailment.
9. `REP-DESCRIPTION` says, “Group A performed 50% better because of the intervention.” It reuses numbers but changes the construct and invents causality.

Each representation has a SHA-256 digest in `manifest.json`. The validator recomputes those hashes. This can detect a changed string, not determine semantic correctness.

The short alt is the compact accessible carrier for the comparison; the long description supplies visual structure and the causal ceiling. The embedding configuration is retained only as transformation identity: an embedding is not a quotation, an accessible description, or claim-level evidence.

## 4. Claim-by-representation diagnosis

Use four dispositions: `supports`, `incomplete`, `contradicts`, and `not_auditable`.

| Representation | A=12 | B=8 | unresolved=4 | no causal basis | Diagnosis |
|---|---|---|---|---|---|
| Pixel evidence declaration | supports | supports | supports | supports | sufficient for fixture claim |
| English alt | supports | supports | supports | incomplete | evidence-bearing but causal ceiling is contextual |
| Chinese alt | supports | supports | supports | incomplete | semantically aligned, human review pending |
| Caption | incomplete | incomplete | absent | incomplete | direction only |
| OCR | supports | absent | supports | incomplete | failed extraction; preserve |
| Long description | supports | supports | supports | supports | accessible core evidence |
| Retrieval record | not_auditable | not_auditable | not_auditable | not_auditable | identity/rank only |
| Generated description | supports | supports | absent | contradicts | adds undefined performance and cause |

The OCR row is **incomplete**, not contradictory: it fails to carry the B value but does not assert another value. The generated description contradicts the evidence ceiling even though its two numbers permit the 50-percent arithmetic. “Because” is the decisive unsupported term. This is why a single semantic-similarity score would be an inadequate audit.

The English alt is concise and task-equivalent for the count comparison. The long description is required for visual structure and the no-causality annotation. If this chart were a core assessment item, the table alternative should be programmatically associated with it. L07 AB-05 reminds us that accessible text existing somewhere is insufficient if the claim-to-source association is lost.

## 5. Multilingual entity, locale, and pooling record

The entity is intentionally generic: `ENTITY-W11-GROUP-A` and `ENTITY-W11-GROUP-B`. “A组” and “B组” are translated labels tied to those identifiers; they are not inferred real-world entities. No transliteration is needed because the labels are symbols, but the ledger still records original script, translated label, locale, and reviewer status.

The English cell is `en-US`; the Chinese cell is `zh-CN`. Both are course-authored strings rather than outputs from a named commercial model. The transformation identity is `human-authored-synthetic-v1`. Time is the frozen package date. Search and tool state are `off`. Geography and account are `none`. The Chinese record is pending bilingual/cultural review, so cross-locale pooling is **blocked**.

That decision is not a judgment that the Chinese sentence is wrong. It means the course has not completed affected-language review or measurement-equivalence assessment. A bilingual reviewer must confirm count terms, unresolved-status wording, task meaning, and any cultural implications. If model-generated descriptions were compared later, each cell would also need provider, surface, observable version, model identity, prompt, repetition, time, account, geography, search state, and tool state. Changing language and locale simultaneously would prevent a language-only causal interpretation.

PAPER-01 cannot fill this gap. Its time provenance conflicts with the course freeze, so W11 permits only a provenance audit and excludes substantive cultural claims.

## 6. Provenance integrity diagnosis

Suppose an S04-compatible manifest validates the asset binding. The permitted conclusion is narrow: under the specified validation procedure, the recorded manifest and asset binding are intact. The following questions remain open and require separate evidence:

- Are the fictional counts truthful about any real population? Not applicable; the case is fictional.
- Who authored an external asset? No inference from integrity alone.
- Does the course have reproduction rights? Established here by course authorship, not by S04.
- Is the evidence accessible and task-equivalent? Requires accessibility review.
- Did a hidden system use pixels, alt, OCR, or caption? Not observable here.
- Did a representation affect ranking? Not tested.

Thus provenance integrity is not truth, authorship, rights, accessibility, or ranking. The manifest helps stabilize an audit object; it does not resolve every audit dimension.

## 7. Disconnected action trace

The action ledger contains three events.

**`EVENT-REC-001` — recommendation.** Text suggests, “Create a draft correction memo explaining that the generated description exceeds the evidence.” No tool name, authorization, attempt, or side effect exists. State: `recommendation_only`.

**`EVENT-PROP-001` — denied proposal.** A structured proposal names `publish_correction` with an external-looking operation class, but its target is `none`, permission is `denied`, and execution is `not_executed`. The proposal is retained to show that structure is not action. There is no tool return and no rollback because no attempt occurred.

**`EVENT-SIM-001` — authorized offline simulation.** A course fixture proposes `create_draft` against `synthetic://w11/draft-buffer`. Permission is scoped to one offline buffer, one operation, and one fixture run. Expected side effect is one synthetic draft record. The simulated return reports a draft identifier. Verification checks the inert local trace, not an external system. A rollback token marks how the synthetic record would be removed, and the rollback state is `available_not_needed`. No network, account, credential, notification, billing event, publication, or user-visible effect exists.

These events support different conclusions. The first supports only that a recommendation was written. The second supports that a proposal was denied and not attempted. The third supports that an authorized offline simulation reached a verified synthetic state. None supports a live agent action or platform outcome. PAPER-39 motivates the permission and containment questions but does not provide universal vulnerability rates or defense guarantees.

## 8. L07 deterministic reproduction

From the workspace root, create a fresh temporary directory and run:

```bash
out_dir=$(mktemp -d /tmp/w11-l07.XXXXXX)
python3 course-labs/L07_drift_adversarial_governance/scripts/review_risk.py \
  --system course-labs/L07_drift_adversarial_governance/data/system_model.json \
  --cases course-labs/L07_drift_adversarial_governance/data/abuse_cases.csv \
  --controls course-labs/L07_drift_adversarial_governance/data/controls.csv \
  --results course-labs/L07_drift_adversarial_governance/data/test_results.csv \
  --governance course-labs/L07_drift_adversarial_governance/data/governance_crosswalk.csv \
  --vocabulary course-labs/shared/fixtures/governance_risk_vocabulary.json \
  --output "$out_dir"
```

Expected stdout is `L07 PASS: 6 cases, 1 escape(s), release BLOCKED`. The output `release_decision.json` records audit status `PASS`, release decision `BLOCKED`, blocking case `AB-05`, three blocked cases, two detected cases, and one escaped case. Its claim ceiling is synthetic control tests and status-vocabulary audit only.

Do not reuse the temporary output as a mutable source. The W11 validator runs the command again, confirms the six frozen input hashes, parses output JSON, checks the decision fields, and compares deterministic output hashes. If any check fails, validation fails closed.

## 9. Source-role audit

PAPER-02 supports a defensive question about coordinated image/text manipulation within evaluated VLM ranking conditions. It does not establish a universal weakness or permission for attack reproduction. PAPER-25 supports a bounded caption-in-context proxy, not proof of direct pixel use. PAPER-39 supports bounded recommendation-agent threat and defense analysis, not a claim about every agent. PAPER-05 is inspected for architecture and evidence-quality gaps only; it supplies no reproduced causal traffic or user-growth effect. PAPER-01 stays quarantined. S04 supports only its technical provenance-integrity ceiling.

The worked case uses no copied figure or prose from any route. Citation of a source does not imply that its entire method or conclusion is absorbed. Every route is paired with an explicit role and boundary.

## 10. Release decision and learner deliverable

The correct decision is **BLOCKED for external release**. The immediate blocking reason is the reproduced L07 AB-05 escaped high-residual accessibility case. The bilingual asset also remains unpoolable pending human review, and the generated description exceeds the evidence ceiling. The offline action simulation itself stays disconnected and does not create an external risk claim.

The learner submits:

1. an `asset-audit.csv` schema with one row per representation and explicit claim dispositions;
2. an original `agent-boundary.svg` specification showing all seven action states, permission gates, logging, and rollback;
3. a locale/model/pooling record with a blocked pooling decision;
4. the L07 command, input hashes, stdout, and parsed decision;
5. a deviation log, including `none observed` when appropriate; and
6. a six-sentence bounded memo covering representation, accessibility, multilingual validity, provenance, action state, and release.

A defensible final sentence is: “Within the frozen course-owned fixture, English alt and the long description carry the count evidence, OCR omits one value, the Chinese row remains pending review and unpooled, and only an authorized disconnected simulation reaches a verified synthetic action state; S04 integrity would not establish truth, rights, accessibility, or ranking, and L07 blocks external release because AB-05 escaped.”
