Intervene

W11 · 5 h

Multilingual, multimodal, and agentic surfaces

Essential question

What changes when the source is an image or the system can act?

What you should be able to do

  1. Trace evidence across modalities.
  2. Identify language and cultural-validity risks.
  3. Extend a threat model to tool-using agents.
PrerequisitesW05 attribution vocabulary.Accessibility basics for text alternatives and captions.
Builds on

W05 attribution vocabulary. · Accessibility basics for text alternatives and captions.

Investigates

What changes when the source is an image or the system can act?

Feeds

Multimodal asset audit feeding L07.

Arrive with a prepared artifact

Reading route
Core PAPER-02/PAPER-25/PAPER-39; PAPER-05 industrial architecture case; PAPER-01 time audit; standards candidate S04.
Viewing route
Open the complete text-first lecture packageSeven-minute modality/provenance boundary explainer. The current equivalent is notes, slide script, worked case, and no-video transcript; no recording is claimed.
Readiness check
Write an evidence-bearing caption and alt text for one supplied figure.
Bring
Accessibility/provenance card.

Concepts, assumptions, and boundary

Evidence can be transformed across modalities

A visual claim may pass through image, alt text, caption, optical character recognition, embedding, retrieval, and generated description before a user sees it. Each transformation can remove context or introduce error. Captions and alt text should communicate the evidence and function of an image, not stuff promotional phrases. Provenance systems can bind origin and editing history, but provenance integrity does not establish truth, authorship, or ranking effect.

Language and agency change the validity boundary

Translation can alter entity names, culturally specific concepts, sentiment, and claim scope. A multilingual result should document locale, language model or translator, validation method, and pooling decision. Agentic systems add tool selection, action execution, permission, and downstream outcome stages. The learner must distinguish a recommendation in text from an executed action and consider authorization, rollback, logging, and harm at every tool boundary.

Inspect the mechanism or evidence structure

MechanismCross-modal evidence and action path
Cross-modal evidence and action pathEvidence moves from source text and image through descriptions and provenance into retrieval. A later branch can remain a recommendation or cross a permission boundary into an action. Failure and rollback points are labeled.Source textImageCaption / altProvenanceRetrievalResponseTool callOutcomepermission boundary
Figure design. Parallel text, image, caption, provenance, retrieval, response, tool-call, and outcome lanes with trust boundaries and accessible equivalents. Recommendation and executed action sit on different sides of permission.
Long description

Evidence moves from source text and image through descriptions and provenance into retrieval. A later branch can remain a recommendation or cross a permission boundary into an action. Failure and rollback points are labeled.

One action, one feedback state

Action

Inspect the same evidence package in text-only, image-caption, and agent-action views.

Feedback

Feedback identifies lost context, unsupported visual claims, inaccessible alternatives, and unauthorized action assumptions.

Accessible alternative

All three views are supplied as static text records with a comparison matrix.

Open the intervention-diff explorer

Produce a reviewable intermediate file

Task
Audit a licensed synthetic multimodal evidence package and extend its threat model.
Inputs
Text, image, caption, provenance manifest example, agent-action trace.
Timebox
70 minutes
Intermediate file
asset-audit.csv and agent-boundary.svg
Stop condition
Reject unverifiable visual claims, inaccessible core evidence, or actions lacking explicit authorization and rollback.

Why each controlled source is here

Complete, retrieve, and revise

Checkpoint
Multimodal asset audit feeding L07.
Reflection
Name one transformation that can change claim meaning.
Revision
Add modality-specific provenance, long description, locale, and permission state.
Low-compute route
Use frozen multimodal files, captions, and action traces; no vision model or live agent is required.

Five retrieval questions

01What does content provenance not prove?

It can protect integrity and editing history, but does not by itself prove truth, authorship, quality, or visibility.

02Why separate recommendation from action?

An executed tool call introduces permissions, side effects, logging, safety, and rollback requirements absent from text alone.

03What must a multilingual comparison record?

Language, locale, translation/model process, entity handling, validation, sampling, and whether results were pooled.

04What is the evidence boundary for W11?

The cases support bounded modality and architecture questions; they are not evidence that a named commercial agent uses the same hidden pipeline.

05What must you submit or revise after this week?

Multimodal asset audit feeding L07. Add modality-specific provenance, long description, locale, and permission state.

Continue in the practice package