# W13 Slide Script — Field Experiments and Monitoring

Production note: every visual below is an original instructional specification. No external figure, screenshot, logo, interface reconstruction, or TeX artwork is reused. Tables and alt text carry the full meaning.

## Slide 01 — Learning while the platform changes

**On-screen text**

Field Experiments and Monitoring. Essential question: How do we learn under a changing platform? Observe → log → validate → decide → preserve → rebaseline or rollback.

**Visual specification**

Create an original circular lifecycle with six labeled stages writing into a central append-only ledger. Add branches for continue, pause, rebaseline, stop, and rollback. No arrow may bypass validation or evidence preservation.

**Speaker notes**

Open with the central difficulty: the measurement surface can change while we observe it. Models, indexes, interfaces, aliases, accounts, locale, source corpora, collection code, and rate controls can move. W13 does not promise to make those variables observable. It teaches how to record known state, preserve unknown state, and make a bounded action under frozen rules. The package uses PAPER-29 and PAPER-38 as bounded core routes, PAPER-12 and PAPER-23 as extensions, PLAT-04 as a checked interface case, and LAB-L06 as the deterministic panel.

**Teaching check**

Ask learners to name one platform-state variable that can change without a visible response-label change.

**Alt text**

A six-stage monitoring loop writes every step to a ledger and branches into five governed decision states.

## Slide 02 — Outcomes and hard boundaries

**On-screen text**

Design blocked observations. Log aliases and explicit configurations. Preserve events and deviations. Plan sequential looks. Detect regime warnings. Apply continue, pause, rebaseline, stop, rollback.

**Visual specification**

Draw seven stacked outcome cards beside a vertical boundary bar. The bar labels “synthetic snapshots,” “procedural reproduction,” and “no live or hidden-backend inference.” Use text and shape, not color alone.

**Speaker notes**

Learners will produce a system card, frozen look schedule, append-only event ledger, incident decision table, and rollback record. The core route has no live access, paid account, browser automation, API call, or network dependency. A deterministic pass establishes the authored panel and decision procedure. It does not establish that the action is correct for every production environment, that a detector works live, or that any platform backend changed.

**Teaching check**

Complete the sentence: “The incident simulator proves the rule is reproducible; it does not prove ___.” Accept production effectiveness or live causality.

**Alt text**

Seven outcomes align with a boundary bar separating synthetic procedural evidence from live-platform and backend claims.

## Slide 03 — The platform belongs in the record

**On-screen text**

Observation = query + response + configuration + time block + session + account + locale + order + failure state + release log + artifact pointer.

**Visual specification**

Create a two-column contrast. The left response-only card contains query, answer, time. The right governed record adds ten labeled state fields and a raw-artifact hash. Mark missing fields “unknown,” never inferred.

**Speaker notes**

A response-only record encourages a simple before/after narrative. The governed record exposes alternative explanations. Account features, locale, order, parser revisions, and known releases can change comparability. `unknown` is a legitimate value when a provider does not reveal backend state. Monitoring does not remove confounding; it tells us when assumptions are unsupported and which action is authorized.

**Teaching check**

Which value is better than an invented model snapshot? `unknown`, paired with the observable product label and time.

**Alt text**

A sparse response record is compared with a full state-aware observation containing configuration, operational, temporal, and evidence fields.

## Slide 04 — Freeze the plan and system card

**On-screen text**

Freeze unit, panel, blocks, metrics, missingness, retries, planned looks, change alerts, action thresholds, budget, authorization, privacy, and rollback target.

**Visual specification**

Design a versioned plan sheet feeding two system cards. Each card shows public label, alias/explicit state, locale, account, session, collection mode, known release, parser, unknown variables, and file hash.

**Speaker notes**

Give the plan an ID, version, timestamp, owner, reviewer, and hash. A later change becomes a deviation event. The system card describes the observable configuration, not an imagined backend. Instructor snapshots add synthetic status and file hashes. Real collection would also require explicit authorization scope, provider terms, budget approval, privacy rules, and credential controls; those are intentionally absent from the core activity.

**Teaching check**

What happens if the retry rule changes after seeing failures? Append a deviation; do not rewrite the original plan.

**Alt text**

A hashed observation plan supplies two detailed system cards, with unobservable fields retained as unknown.

## Slide 05 — Dynamic alias versus explicit configuration

**On-screen text**

Alias: convenient, potentially remapped. Explicit configuration: more traceable, still incomplete. Record both; neither reveals every backend variable.

**Visual specification**

Draw two labeled lanes. The alias lane enters an opaque routing box with several possible exits. The explicit lane enters a named configuration box but still passes through opaque index, cache, and policy boxes. Use question-mark labels for unknowns.

**Speaker notes**

A provider can update what a dynamic alias points to without changing the alias. An explicit label can improve longitudinal identity but may leave routing, corpus, search state, and tool policies hidden. Comparing routes in the same time block controls some observable timing differences. It does not reverse-engineer the pipeline. The system card must preserve exactly which fields were visible and which were not.

**Teaching check**

Can an explicit model label prove the retrieval index version? No, unless that index identity is separately documented.

**Alt text**

Alias and explicit routes both pass through partly opaque state, although the explicit route provides a more stable recorded label.

## Slide 06 — Same-time difference, unknown mechanism

**On-screen text**

B-02 mention: SYN-A 25/59 = 0.423729; SYN-B 31/60 = 0.516667; difference 0.092938. Descriptive, synthetic, mechanism unknown.

**Visual specification**

Create two adjacent fraction bars with exact numerators, denominators, and rates. Connect them to a difference bracket. Under the bracket, place three labeled unknown boxes: routing, corpus/index, authored fixture design.

**Speaker notes**

The package maps SYN-A to an authored alias-route card and SYN-B to an authored explicit-route card. This is teaching metadata. The same-time difference is real inside the frozen CSV arithmetic. A same-time alias difference does not identify a hidden backend change. There is no production backend in the fixture. Even in real collection, route label, retrieval state, session, and model behavior could contribute.

**Teaching check**

Rewrite “the alias used a worse model” as a valid statement. Expected: “The recorded routes differed by 0.092938 under the B-02 synthetic panel; cause is not identified.”

**Alt text**

Exact same-block mention rates differ by 0.092938, while three unknown-mechanism boxes prevent causal backend attribution.

## Slide 07 — Time blocks are comparability windows

**On-screen text**

A block closes on a planned batch, known release, configuration change, parser revision, outage, or protocol deviation. Preserve every regime.

**Visual specification**

Draw a timeline with B-01, B-02, and B-03 as containers. Place a release marker at the B-02 boundary and an outage inside B-03. Use vertical bars, dates, and written trigger labels.

**Speaker notes**

A time block is not just a calendar label. The plan explains why observations inside it are treated as comparable. A known configuration change can close a block even if outcomes look stable. An unexplained metric discontinuity can trigger investigation without naming a cause. Rebaseline preserves old and new regimes rather than deleting the earlier period.

**Teaching check**

Does rebaseline mean erase B-01? No. Archive it and prohibit unqualified pooling across the boundary.

**Alt text**

Three comparability blocks are separated by a known release and interrupted by a later outage, with every regime retained.

## Slide 08 — Session, account, locale, and privacy

**On-screen text**

Record what was cleared, what account was used, which locale applied, and what remains unknown. Never use unapproved credentials or personal data.

**Visual specification**

Create a four-quadrant system card for session, account, locale, and privacy. Each quadrant has observed, controlled, unknown, and prohibited fields. Provide a matching text table.

**Speaker notes**

Prior turns, cookies, cache, personalization, subscriptions, and region settings can matter. “Fresh session” needs an operational definition. A shared account can contaminate state and create privacy risk. Locale includes language, region, and sometimes location. The supplied LAB-L06 route avoids these live dependencies. Optional collection would need a separate approval and data-handling plan.

**Teaching check**

Is access to a web interface equivalent to authorization for automation? No. Method and scope require explicit review.

**Alt text**

A four-part card distinguishes observed, controlled, unknown, and prohibited session, account, locale, and privacy states.

## Slide 09 — Collection order and failures are data

**On-screen text**

Counterbalance order. Preserve realized sequence. Classify timeout, refusal, parser error, unavailable panel, authorization denial, and rate-limit response separately.

**Visual specification**

Show a counterbalanced A–B/B–A order grid beside six distinct failure-state cards. An interruption arrow preserves completed cells and leaves unobserved cells explicitly marked, never silently removed.

**Speaker notes**

Collecting every SYN-A cell before every SYN-B cell can confound surface with time. Counterbalancing reduces systematic order effects but cannot freeze the platform. Preserve the actual order and duration. Failures have different meanings and denominator consequences. A retry needs a frozen reason, maximum count, delay, and terminal state. Never retry until a preferred answer appears.

**Teaching check**

Why should an authorization denial not be coded as a zero mention? It is an access-state incident, not an observed answer outcome.

**Alt text**

A counterbalanced collection grid and six typed failure cards show that order and operational failures remain explicit records.

## Slide 10 — Known releases and unknown change alerts

**On-screen text**

Known state event: documented release or configuration revision. Outcome alert: observed shift under a frozen threshold. Neither alone proves outcome cause.

**Visual specification**

Create two parallel change-point lanes. The top begins with an official event card; the bottom begins with a metric-threshold alert. Both converge on investigation, but only the known-state lane carries an identified configuration change label.

**Speaker notes**

Keep release publication, effective, and retrieval dates distinct. A documented configuration change can justify rebaseline because the design object changed. An outcome alert such as a rate shift or missingness spike says the regime may have changed. It does not say why. Even when a release and shift coincide, causal attribution needs more design and assumptions.

**Teaching check**

Can a known release be logged when no outcome moves? Yes; comparability can still change.

**Alt text**

Known-state and outcome-alert lanes both trigger investigation, but only the former documents a configuration event and neither proves causality.

## Slide 11 — Append-only event ledger

**On-screen text**

Sequence, event ID, occurred/recorded time, type, block, system card, facts, evidence, deviation, decision, reason, parent, previous hash, row hash.

**Visual specification**

Draw eight ledger rows as linked cards. Each row points to the previous row hash and a parent event. Show correction as a new row pointing backward rather than erasing an earlier row.

**Speaker notes**

The ledger links plan freeze, block completion, releases, deviations, incidents, decisions, rollback, and closure. A canonical sorted-key JSON serialization generates each SHA-256 row hash. The first previous hash is 64 zeroes; later rows use the preceding hash. This makes silent mutation detectable under the file process. It does not guarantee truth, secure storage, or trustworthy clocks.

**Teaching check**

How do you fix an incorrect event description? Append a correction event with the original ID; do not overwrite the row.

**Alt text**

Eight sequential ledger cards form a hash chain, while a correction adds a new linked record instead of mutating history.

## Slide 12 — Planned event, deviation, or incident?

**On-screen text**

Planned event follows protocol. Deviation departs from plan. Incident threatens authorization, privacy, cost, rate limits, policy, evidence, or measurement integrity.

**Visual specification**

Create a three-column classifier with example cards: scheduled block look, extra retry, parser hotfix, outage, authorization expiry, privacy exposure. Arrows show that a deviation can escalate into an incident.

**Speaker notes**

Precise event typing supports the right response. A scheduled look is not a deviation. An unplanned outcome peek is. A temporary outage is an incident because it threatens completeness and order. A privacy event is a hard-gate incident. The same row can record both a deviation code and incident severity when the plan departure creates risk.

**Teaching check**

Classify a parser version changed mid-block without approval. It is a deviation and may trigger a measurement-integrity incident.

**Alt text**

Three columns distinguish planned, deviating, and incident events, with examples and an escalation path.

## Slide 13 — Plan sequential outcome looks

**On-screen text**

Outcome looks: after complete B-01, B-02, B-03 only. Hard-gate monitoring: continuous. Unplanned peek: log deviation and lower claim ceiling.

**Visual specification**

Use a timeline with three closed-eye/open-eye analysis markers at block ends. Place a continuous guardrail rail below for authorization, privacy, cost, policy, and evidence preservation.

**Speaker notes**

Repeated outcome peeking creates researcher freedom. W13 freezes three block-end looks and makes no efficacy stopping claim. Operational hard gates are monitored as events arrive; a privacy or authorization failure cannot wait. Continuous guardrail monitoring does not authorize continuous analysis. A formal sequential inference plan would require information times, boundaries, multiplicity control, and delayed-data rules that are not invented here.

**Teaching check**

Why can privacy be checked continuously while mention rate is not repeatedly interpreted? One is a hard gate; the other is an outcome subject to analysis freedom.

**Alt text**

Three scheduled outcome looks sit above a continuous operational guardrail rail, emphasizing their different purposes.

## Slide 14 — Change-point alerts do not name causes

**On-screen text**

Alert if absolute block shift ≥ 0.10 or noncomplete proportion > 0.10. Investigate regime. Do not announce a platform cause.

**Visual specification**

Draw a simple rate timeline crossing a horizontal 0.10 alert band. From the alert, branch to possible explanations: release, order, query mix, missingness, parser, fixture, chance. Label all as candidates.

**Speaker notes**

The numeric threshold is an authored teaching rule, not a validated production detector. A short three-block series cannot estimate false alarms or tune a change-point method. The alert initiates evidence preservation and review. A known configuration event can independently trigger rebaseline. An unexplained rate shift remains an unknown-cause regime warning.

**Teaching check**

What does crossing 0.10 establish? Only that the frozen alert rule fired, not that the model or index changed.

**Alt text**

A rate crosses a declared alert threshold and branches to multiple possible causes, none selected by the alert alone.

## Slide 15 — Five decision states

**On-screen text**

Continue: protocol valid. Pause: recoverable issue. Rebaseline: comparability changed. Stop: hard gate failed. Rollback: restore or contain an approved prior state.

**Visual specification**

Create five equally sized decision cards with definition, owner, expiry or closure condition, and required evidence. Avoid arranging them as a simple severity ladder because rebaseline and pause can coexist.

**Speaker notes**

Continue is not success; it means the authorized design remains valid. Pause preserves state while diagnosing. Rebaseline closes one regime and starts another. Stop ends the collection path. Rollback is a tested action, often attached to stop or containment. Every action writes an event and names the next authorized step.

**Teaching check**

Can rebaseline occur when metrics are stable? Yes, if the configuration identity or protocol changed.

**Alt text**

Five decision cards define distinct monitoring actions and their evidence requirements without implying a single severity order.

## Slide 16 — Decision precedence

**On-screen text**

Stop/required rollback > pause > rebaseline > continue. Authorization, privacy, evidence preservation, policy integrity, and cost cap override favorable outcomes.

**Visual specification**

Build a top-down decision tree. Hard-gate failure leads to stop and optional required rollback. Otherwise recoverable operational failure leads to pause; otherwise comparability change leads to rebaseline; otherwise continue.

**Speaker notes**

Precedence prevents post hoc convenience. A favorable mention rate cannot override expired authorization. A pause can coexist with a pending rebaseline, but collection remains paused until resume criteria are met. A rollback target includes version, artifacts, authorizing role, verification, and follow-up. Overrides for safety are allowed but become explicit ledger events.

**Teaching check**

If authorization expired and the metric improved, what is the decision? Stop; outcome direction is irrelevant to the hard gate.

**Alt text**

A decision tree gives hard-gate stop and rollback priority over pause, rebaseline, and continue.

## Slide 17 — LAB-L06 frozen panel

**On-screen text**

20 queries × 2 surfaces × 3 blocks × 3 repetitions = 360 events. 354 complete. Six noncomplete. Five distinct outcome fields.

**Visual specification**

Create a four-dimensional panel cube expanded into a table. Label query, surface, block, and repetition axes. Place six written missing-state markers and list mention, citation, entailment, absorption, referral separately.

**Speaker notes**

LAB-L06 is synthetic, offline, deterministic, and standard-library only. Every designed cell exists, including noncomplete rows. Outcomes are scored only among complete responses. Repeated rows within one query are dependent. Wilson intervals are descriptive and not dependence-adjusted. The panel cannot establish a live cross-engine or causal effect.

**Teaching check**

Why must six noncomplete rows remain? Their absence would silently change the designed panel and obscure operational state.

**Alt text**

A panel cube represents 360 designed query, surface, block, and repetition cells, including six explicit noncomplete events.

## Slide 18 — Descriptive drift is not mechanism

**On-screen text**

SYN-A mention: 0.482759 → 0.400000, delta −0.082759. SYN-B: 0.534483 → 0.491525, delta −0.042958. Synthetic drift only.

**Visual specification**

Plot two labeled three-point lines with exact rates and numeric first-to-last deltas. Use different line styles and point shapes. Add a broad caption: “No mechanism identified.”

**Speaker notes**

The analyzer reports first-to-last descriptive differences. It does not know about product releases, interventions, routing, or demand. Block comparisons can motivate monitoring questions, but they cannot identify cause. Missingness differs slightly by surface and block, and repeated-query dependence remains. Preserve event-specific metrics; do not collapse mention with citation or absorption.

**Teaching check**

May we call the SYN-A decline a model update effect? No. It is an authored fixture pattern with no identified mechanism.

**Alt text**

Two synthetic mention-rate lines decline by different amounts while an explicit label states that no mechanism is identified.

## Slide 19 — Incident one: continue

**On-screen text**

INC-01: B-01 has 4/120 noncomplete = 0.033333. Continue threshold ≤ 0.05; hard gates pass. Decision: CONTINUE.

**Visual specification**

Show a fraction card beside a 0.05 threshold gauge with exact values. A gate checklist confirms authorization, privacy, policy, cost, and evidence preservation all pass.

**Speaker notes**

The decision follows a frozen quality rule. Continue means the next planned step may proceed; it does not mean the outcome is favorable or the panel is repeatable in production. The four noncomplete events remain in the dataset. If the threshold had been chosen after observing 0.033333, the rule would be post hoc.

**Teaching check**

What evidence would invalidate continue despite low missingness? Any hard-gate failure, such as expired authorization or privacy exposure.

**Alt text**

A 0.033333 noncomplete rate sits below a 0.05 threshold while all five hard gates pass, producing continue.

## Slide 20 — Incident two: rebaseline

**On-screen text**

INC-02: authored configuration record changes before B-02; alias backend remains unknown. Decision: REBASELINE. Preserve B-01 and begin a new regime.

**Visual specification**

Draw a vertical regime boundary before B-02. Place the changed explicit configuration card on the boundary and an alias card labeled “backend unknown.” Separate pre- and post-boundary summary boxes.

**Speaker notes**

Rebaseline is triggered by a known configuration-record change, not by the direction of the outcome. The alias still lacks a stable backend identity. The B-02 same-time route difference remains descriptive. Do not claim the changed record caused that difference. Archive B-01, start the new regime, and prohibit unqualified pooling across the boundary.

**Teaching check**

What is known? The authored configuration record changed. What remains unknown? Any hidden backend mapping or causal role.

**Alt text**

A recorded configuration change creates a new regime at B-02 while the dynamic alias backend remains explicitly unknown.

## Slide 21 — Incident three: pause

**On-screen text**

INC-03: six unavailable cells in a 20-cell mini-batch = 0.30, above 0.10 pause threshold. Decision: PAUSE and preserve order.

**Visual specification**

Create a 20-cell grid with six cells labeled unavailable. A pause branch lists evidence preservation, owner, diagnosis, expiry, and resume criteria. Do not depict retries occurring automatically.

**Speaker notes**

The simulated outage is potentially recoverable, and authorization and privacy remain valid. Pause protects order and evidence while the team investigates. Resume requires a new ledger event and satisfied criteria. The interrupted mini-batch is an overlay exercise; it does not modify LAB-L06 or fabricate new panel outcomes.

**Teaching check**

Why not retry every unavailable cell immediately? Retries need a frozen rule, rate-limit compliance, preserved first attempts, and a valid regime.

**Alt text**

Six of twenty mini-batch cells are unavailable, crossing the pause threshold and activating a documented recovery path.

## Slide 22 — Incident four: stop and rollback

**On-screen text**

INC-04: authorization expired; projected cost 12 exceeds remaining cap 8. Decision: STOP_AND_ROLLBACK to approved snapshot-only mode.

**Visual specification**

Draw two hard-gate cards, authorization and budget, both failing. A stop arrow leads to a rollback manifest containing target state, affected artifacts, authorizer, verification hash, and follow-up review.

**Speaker notes**

One hard-gate failure is sufficient. Here two fail. No live request occurs, and unapproved rows are not admitted. Rollback restores the approved course snapshot mode and verifies the target. It does not erase the incident. The ledger retains expiry, projected cost, decision, actor, and rollback completion.

**Teaching check**

Could the team proceed because only four more calls are needed? No. Scope and cap are authorization boundaries, not negotiable convenience.

**Alt text**

Expired authorization and an exceeded projected budget lead to stop and a verified rollback to the approved offline snapshot state.

## Slide 23 — Reading and interface boundaries

**On-screen text**

PAPER-29: study-specific repeats. PAPER-38: vendor observational data. PAPER-12: fixed slate, v2. PAPER-23: reconstructed pipeline, v2/KDD 2026. PLAT-04: bounded preview labels.

**Visual specification**

Create a five-row evidence card table with public ID, admitted role, status/version, and prohibited inference. Crosshatch commercial-unreleased data and label all artifact/reproduction limits in text.

**Speaker notes**

PAPER-29 does not provide universal run counts or isolate token sampling. PAPER-38 is a single-author vendor preprint with unreleased commercial data; proposed protocols are not completed findings. PAPER-12 conditions on a fixed candidate slate and does not show live retrieval or business impact. PAPER-23 measures a reconstructed pipeline, not a disclosed commercial engine. PLAT-04 documents specified Microsoft interface counts; those are not rank, authority, placement, or cross-engine share.

**Teaching check**

Which source licenses us to infer a hidden backend from an alias difference? None.

**Alt text**

Five evidence cards pair each public identifier with its admitted use, version or status, and a prohibited generalization.

## Slide 24 — Exit test and bounded decision

**On-screen text**

State recorded? Looks planned? Cause known or unknown? Threshold frozen? Hard gates honored? Ledger append-only? Rollback verified? Backend inference avoided?

**Visual specification**

Build an eight-gate exit path ending in a versioned monitoring package: plan, system cards, panel, ledger, incident table, decision trace, rollback record, and boundary statement.

**Speaker notes**

Close with the deterministic decisions: INC-01 continue, INC-02 rebaseline, INC-03 pause, INC-04 stop and rollback. Another reviewer must reproduce them from the same inputs and precedence. That proves procedural consistency only. No W13 recording exists, no timed pilot has occurred, and no live automation is authorized. The strongest unknown is the hidden system state behind observable labels.

**Teaching check**

Final prompt: “What changed, what evidence records it, what remains unknown, and which action is authorized next?”

**Alt text**

Eight reproducibility gates lead to a versioned monitoring package with explicit decisions, evidence, rollback, and hidden-backend uncertainty.
