# W13 No-Video Equivalent — Field Experiments and Monitoring

## Status and use

No W13 recording exists. This document is a planned no-video equivalent and a possible future narration script. It is not a transcript of delivered instruction. The chapter durations are instructional planning budgets, not observed timestamps or classroom measurements. Any later recording requires verified duration, synchronized captions, speaker labels, description of figures and ledgers, and navigable chapters.

The route uses frozen instructor-provided snapshots. It performs no live platform collection, browser automation, API call, account access, purchase, or personal-data processing.

## Planned chapter budget

| Chapter | Focus | Planned minutes |
|---:|---|---:|
| 1 | The platform as part of the time series | 4 |
| 2 | Observation plans and system cards | 4 |
| 3 | Dynamic aliases and explicit configurations | 4 |
| 4 | Blocks, order, failures, and releases | 4 |
| 5 | Append-only events and deviations | 5 |
| 6 | Sequential looks and change points | 4 |
| 7 | Decision states and hard gates | 4 |
| 8 | LAB-L06 reproduction | 4 |
| 9 | Four synthetic incidents | 5 |
| 10 | Reading boundaries and conclusion | 4 |

**Total planned route: 42 minutes.** The design range is 35–45 minutes, pending rehearsal.

## Chapter 1

How do we learn under a changing platform? The first answer is to stop treating the platform as a fixed backdrop. A field observation happens while model labels, routing, indexes, source corpora, interfaces, account features, search settings, collection code, rate controls, and policies may change. A response and a timestamp cannot tell us which of those states remained comparable.

Suppose a visibility rate differs between June and July. We could tell a simple story about an intervention. But query composition may have moved. A product alias may point somewhere new. An interface may expose citations differently. A parser may have changed. A collection outage may have affected one query stratum. An account or locale may differ. None of these explanations becomes true merely because we list it. Monitoring records enough state to see which assumptions are plausible and which are broken.

The minimum record therefore joins the query ID, response artifact, surface, configuration label, alias or explicit state, time block, repetition, captured time and offset, session, account, locale, collection mode, realized order, response status, parser, known release, and raw-artifact pointer. When the provider does not disclose a backend variable, write `unknown`. A plausible model name inferred from output style is not a measurement.

LAB-L06 gives us a controlled teaching panel. It has 20 queries, two synthetic surfaces, three blocks, and three repetitions, for 360 designed events. Six are explicitly noncomplete. The panel lets us distinguish repeated rows within a block from differences between blocks. It does not make time random, reveal a platform state, or identify why a rate changes.

Monitoring is part of identification because it preserves alternative explanations. A release, parser revision, account change, order interruption, or protocol deviation can close a comparability window. The correct decision may be to continue, pause, rebaseline, stop, or rollback. Those are protocol actions, not labels for good or bad outcomes.

The W13 core route uses supplied snapshots and a deterministic incident simulator. No learner must create an account or collect live answers. This matters ethically and scientifically. It removes hidden financial, privacy, authorization, and terms-of-use dependencies while still teaching the full decision logic.

There is another reason to use snapshots: everyone analyzes the same evidence. A live class can otherwise receive different interface states, outages, account entitlements, and costs, making learning depend on access rather than reasoning. The supplied route creates equal inputs and an auditable answer key. It does not claim ecological equivalence to a production platform. The difference between pedagogical control and field realism should be visible, not treated as a defect to hide.

At the end of each block, ask three questions in order. First, is the data package structurally valid? Second, did any hard gate or comparability condition change? Third, what outcome pattern is descriptively visible? Reversing that order invites a favorable pattern to influence quality and safety judgments.

## Chapter 2

The observation plan must be frozen before outcomes are interpreted. It records the experimental or observational unit, query frame, surfaces, configurations, blocks, repetitions, primary metric, guardrails, missing-state policy, retry rule, planned outcome looks, operational checks, change alerts, action thresholds, budget, authorization, privacy handling, and rollback target. Give the plan a version, owner, reviewer, timestamp, and hash.

If the team later changes a threshold, retry count, query list, parser, or collection order, that change becomes a deviation. It may be a justified change. The point is that it remains visible. Editing the original plan after seeing outcomes destroys the distinction between prior rule and later response.

Each observed configuration receives a system card. Record the public product or fixture label, route type, dynamic alias or explicit identifier, version if exposed, interface or API route, account and subscription state, session treatment, search or browsing state, language and region, collection method, known release date, parser, and unobservable fields. An instructor snapshot additionally records `synthetic=true` and file hashes.

Session state needs an operational description. “Fresh session” might mean a new browser profile, cleared conversation history, cleared cookies, or a new API conversation. Those actions are not equivalent, and none guarantees that provider-side state is absent. Account state can control subscription features or personalization. Locale can include language, region, location, and interface setting. Write what is observed and controlled.

Collection order is also a field. If every surface A observation happens before every surface B observation, a surface contrast is also a time contrast. Counterbalancing or randomizing the order can reduce systematic confounding, but it cannot freeze the system. Preserve the realized order, start and end times, interruptions, and whether a retry was allowed.

Failures stay in the dataset or event ledger with typed states. Timeout, explicit refusal, parser error, unavailable citation panel, authorization denial, rate-limit response, and privacy stop are not interchangeable. A missing response is not a zero mention. A retry cannot continue until the preferred answer appears. The plan states an allowed reason, maximum, delay, and terminal status, and every attempt remains linked.

## Chapter 3

Dynamic product aliases and explicit configurations solve different operational problems. A dynamic alias is convenient because it follows a provider’s current offering. That convenience weakens longitudinal identity. The provider can change routing, snapshot, tools, safety behavior, index, or interface without changing the alias.

An explicit configuration improves traceability if one is exposed. Yet a model or version label may still leave backend routing, source index, cache, retrieval policy, system prompt, and tool state unobservable. “Explicit” does not mean “complete.”

Record alias and explicit routes in separate system cards. When permitted, a same-time comparison can hold query, locale, account, session, order, and collection method as close as possible. It documents output differences under the recorded labels. It does not reveal which hidden component caused them.

The W13 fixture maps LAB-L06 SYN-A to an authored alias-route card and SYN-B to an authored explicit-route card. No real product is represented. In B-02, SYN-A mention is 25 among 59 complete responses, or 0.423729. SYN-B mention is 31 among 60, or 0.516667. The difference is 0.092938.

A same-time alias difference does not identify a hidden backend change. In this synthetic case, there is no production backend to infer. In a live observation, routing, index state, cache, search behavior, account state, stochastic generation, parser behavior, and query interactions could all contribute. The correct conclusion is that the recorded routes differ descriptively in the named block.

If a provider publishes an official alias-mapping change, record issuer, announcement, effective date, retrieved date, and exact affected label. That is evidence that a published configuration statement changed. It can justify rebaseline because the recorded design object changed. It still does not prove that a concurrent outcome shift was caused by the mapping.

PAPER-29 is useful precisely because its repeated web-interface study retains this limitation. Same-window differences cannot be assigned solely to token sampling when retrieval, routing, cache, index, and product state are not pinned. Its counts and windows are study-specific, not universal rules. PAPER-38 also cannot support black-box inference about training, ingestion, memory, or retrieval from output patterns.

## Chapter 4

A time block is a declared comparability window. It has an opening condition, closing condition, and explanation for treating observations inside it as comparable. A planned batch can close a block. So can a known release, configuration revision, parser change, account-state change, locale change, outage, or protocol deviation.

Separate known-state events from outcome alerts. A known-state event has a cited release, controlled fixture revision, or approved collection-code change. An outcome alert fires because a rate, missingness level, or quality measure crosses a prespecified threshold. It says the regime deserves investigation. It does not name a cause.

Release dates also have roles. Announcement, effective, observed, and retrieved dates can differ. An interface announcement may precede rollout. The course cannot verify hidden deployment time, so it records the public event and its limitations.

Order matters inside blocks. Counterbalance routes across queries and repetitions when the design allows. If an outage occurs halfway through, the first and second halves may contain different intent strata. Preserve realized order and interruption. Do not delete partial batches or restart until a preferred complete set appears.

Account and session changes can close comparability even when the surface label remains constant. Locale drift can change evidence and answer language. A parser hotfix can change which citations are detected. The monitoring record should make each event visible so the analyst can decide whether to separate regimes.

The W13 change-point teaching rule has two paths. A known configuration or protocol identity change triggers rebaseline directly. A descriptive alert fires if the absolute block-to-block primary-rate shift is at least 0.10 or an incident-batch noncomplete fraction exceeds 0.10. That threshold is authored for teaching. Three blocks cannot validate detector sensitivity or false-alarm performance.

Rebaseline closes the old regime and opens a new one. It does not erase the old data. Report regimes separately unless a justified bridge analysis states its assumptions. Rebaseline can occur even when the metric does not move, because comparability concerns the design state, not outcome direction.

## Chapter 5

The append-only event and deviation ledger is the spine of monitoring. It links the frozen plan, block completion, system-card changes, release notices, observations, failures, deviations, incidents, decisions, rollback, and closure.

Append-only means a prior row is not overwritten. If an event description is wrong, append a correction that names the original event and supplies the revised fact. The history remains visible. This prevents a clean-looking final file from hiding when the team learned about a release or changed a threshold.

A minimum event row contains event ID, contiguous sequence, occurred and recorded times, event type, block, system-card ID, actor role, severity, facts, evidence pointer, deviation code, decision, reason code, parent event ID, previous hash, and row hash. Personal data, credentials, and secrets do not belong in this ledger. Use controlled evidence pointers when sensitive records must be preserved elsewhere.

Distinguish planned events, deviations, and incidents. A scheduled B-01 analysis look is planned. An extra outcome peek or parser change is a deviation. An outage, authorization expiry, privacy exposure, cost-cap threat, rate-limit conflict, or evidence-preservation failure is an incident. A deviation can escalate when it threatens integrity or safety.

The fixture uses a simple SHA-256 chain. Serialize each event without its `row_hash` as compact JSON with sorted keys. The first `previous_hash` is 64 zeroes. Each later `previous_hash` equals the prior row hash. The validator also checks unique IDs, contiguous sequence, monotonic recorded time, allowed event types, valid parent references, and decision precedence.

This chain makes mutation detectable under the chosen canonicalization and file process. It does not prove that the event facts were true, that timestamps came from a secure clock, or that storage cannot be replaced wholesale. Independent review and access control remain separate.

The W13 fixture contains eight rows: plan freeze; B-01 look and continue; configuration change and rebaseline; B-02 look inside the new regime; outage and pause; authorization/cost failure and stop-and-rollback; rollback verification; and evidence-preserving closure. Another reviewer can reconstruct the exact chain without network access.

Event time and knowledge time can differ. Suppose an interface release became effective during B-01 but the team discovered the notice after B-02. The release should not be inserted retroactively between old rows. Append a new discovery event at the current sequence, preserve the cited effective date inside its facts, and reopen affected decisions according to the plan. That record shows what happened, when the team learned it, and what action followed. A single timestamp cannot represent all three.

Every deviation also needs a disposition. `open` means investigation continues. `accepted-with-boundary` means the analysis can continue with narrowed interpretation. `requires-reanalysis` identifies affected outputs. `invalidates-regime` closes comparability. `closed` records reviewer acceptance and next monitoring action. A list of deviations without analytic consequences is not enough.

The ledger should avoid overcollection. It needs evidence pointers and stable identifiers, not full credentials, personal prompts, confidential outputs, or secret tokens. Integrity and privacy can conflict if a team indiscriminately logs everything. The monitoring plan states what can be retained, where, for how long, and who may access it.

## Chapter 6

Sequential outcome looks need a plan because repeated peeking creates researcher freedom. W13 permits one outcome look after each complete block: B-01, B-02, and B-03. It does not define an efficacy-success boundary. The final interpretation remains descriptive.

Operational hard-gate monitoring is different. Authorization, privacy, policy, evidence preservation, rate-limit status, and budget risk must be checked continuously. A privacy event cannot wait until the next outcome look. Continuous safety monitoring, however, does not license continuous interpretation of mention or citation rates.

If an unplanned outcome look occurs, append a deviation. Do not pretend it did not happen. Subsequent confirmatory language may need to become exploratory. A formal sequential inference plan would specify an estimand, information times, decision boundaries, multiplicity treatment, delayed data, missingness, and possibly alpha spending. W13 does not fabricate such a result for the synthetic panel.

Change-point alerts share this discipline. The alert threshold is frozen in advance. Crossing it starts investigation. It does not prove a product update. Possible explanations include a known release, alias mapping, query composition, order, missingness, parser behavior, source-corpus movement, authored fixture structure, or chance variation. If no cause is observed, write `cause: unknown`.

A known release can close a regime without an outcome alert. An unexplained alert can trigger pause without a known release. When a release and shift coincide, they still do not establish causality without a design and assumptions that separate concurrent changes.

Negative controls can strengthen diagnosis when frozen in advance. A parser-only count, collection latency, or fixture field known not to respond to a content intervention may reveal an operational discontinuity. It does not automatically identify the cause. Its expected behavior, denominator, alert threshold, and missing-state policy must be declared before the incident. Choosing a convenient control after seeing the shift creates another degree of freedom.

Response latency is part of monitoring effectiveness. Record when the alert fired, when evidence was preserved, when the owner acknowledged it, and when pause or rollback completed. A mathematically correct threshold is not an effective control if no one responds. The deterministic W13 simulator makes decisions immediately for reproducibility; it does not measure how a real team performs under pressure.

PAPER-38 is especially important to read with this separation. It reports observational repeated tracking from a single-author vendor preprint, unreleased commercial data, provider-specific policies, varying cohorts, and vendor-defined rules. Its proposed protocols remain future designs. They cannot be converted into completed sequential evidence. PAPER-29 motivates repeated observation but retains its study population, exclusions, and unpinned state.

## Chapter 7

The five action states need definitions and precedence.

Continue means the current authorized protocol and regime remain valid, quality stays within threshold, and the next planned step may proceed. It does not mean the primary outcome is favorable.

Pause means temporarily suspend collection while preserving state and investigating a potentially recoverable issue. Examples include an outage, noncomplete spike, rate-limit warning, parser anomaly, approaching budget threshold, or unknown configuration state. Pause has an owner, expiry, and resume criteria.

Rebaseline means a documented configuration or protocol change breaks comparability. Close the prior regime, retain its data, open a new baseline, and avoid unqualified pooling. Rebaseline can be pending while collection is paused.

Stop means end the authorized collection path because a hard gate failed or recovery criteria were not met. Authorization expiry, privacy exposure, prohibited automation, evidence-preservation failure, policy conflict, persistent rate-limit conflict, and cost-cap breach are stop conditions.

Rollback restores or contains a specific approved prior state. It names target version, affected artifacts, actor, verification, and follow-up. Rollback can accompany stop. It does not erase the incident.

The precedence is stop and required rollback first, then pause, then rebaseline, then continue. A favorable outcome cannot override a hard gate. If authorization has expired and the metric looks better, stop. If an outage and configuration change occur together without a hard-gate failure, pause while preserving that a rebaseline will be required. If only configuration identity changed, rebaseline. If no gate, operational, or comparability rule fires, continue.

Every nonterminal action needs an expiry or review time. A pause without expiry can become an undeclared stop. A rebaseline without a minimum new-baseline requirement can invite immediate comparison across unstable regimes. Continue should name the next planned look. Stop should name evidence-preservation and closure responsibilities. Rollback should name a verification that can fail.

Cost and rate limits belong in the plan. Record total cap, warning fraction, expected block cost, remaining amount, approver, and action. Do not evade provider limits through distributed or concealed automation. The supplied route avoids this entirely.

Authorization records method, account, time, scope, data, and revocation. Access to an interface does not automatically authorize automation. Privacy review covers prompts, responses, account metadata, and logs. Unexpected personal data triggers stop and the approved incident procedure. Never place credentials or personal data into the public event ledger.

## Chapter 8

Let us reproduce LAB-L06 before applying incidents. The command reads `panel.csv`, requires at least 20 queries, and writes to a fresh output directory. The expected pass line reports 360 events, 20 queries, two surfaces, and three blocks.

The output contract contains eight files: normalized panel, metric estimates, time drift, query-mix sensitivity, missingness, data dictionary, validity boundary, and run manifest. The validity record reports 354 complete and six noncomplete events. The run manifest marks the route deterministic and offline.

Mention, citation, entailment, absorption, and referral remain separate outcomes. Their denominators contain complete responses only. Noncomplete designed cells remain rows. The Wilson intervals are descriptive and do not adjust for dependence among repeated observations for a query.

For mention, SYN-A rates are 0.482759 in B-01, 0.423729 in B-02, and 0.400000 in B-03. First-to-last difference is minus 0.082759. SYN-B rates are 0.534483, 0.516667, and 0.491525. Its difference is minus 0.042958.

These are authored patterns. They do not establish that one surface drifted because of a model, index, routing, or release change. The same-time B-02 difference of 0.092938 is also descriptive. The external W13 configuration and incident cards are an overlay; they never rewrite the LAB-L06 CSV.

PAPER-12 and PAPER-23 help explain why environment boundaries matter. PAPER-12’s locked v2 testbed conditions on a fixed candidate slate and named model calls. It does not measure live retrieval into that slate or future-version stability. PAPER-23’s locked v2/KDD 2026 pipeline is reconstructed and conditions target selection. It supports stage-specific analysis in that environment, not disclosure of a commercial architecture.

PLAT-04 documents a Microsoft/Bing public-preview interface with citation, cited-page, sampled-grounding-query, and trend labels on specified surfaces. Those labels are authoritative only for the checked documentation. Their counts are not rank, authority, placement, cross-engine share, or business effect.

## Chapter 9

Now apply four incidents using the frozen action rules.

INC-01 occurs after B-01. Across the two surfaces, four of 120 designed events are noncomplete. Four divided by 120 is 0.033333. The continue threshold allows a noncomplete fraction at or below 0.05 when every hard gate passes. Authorization, privacy, policy, evidence preservation, and budget pass. Decision: continue. This says the next planned block is permitted, not that the outcome is good.

INC-02 occurs before B-02. An authored release card changes the explicit configuration record from `SYN-EXPLICIT-2026-06-A` to `SYN-EXPLICIT-2026-06-B`. The dynamic alias continues to expose no stable backend identity. No hard gate or outage applies. Decision: rebaseline. Preserve B-01 as the first regime and open the new regime at B-02. Do not infer that the configuration record caused the B-02 route difference.

INC-03 injects a synthetic outage into a 20-cell B-03 mini-batch. Six become unavailable. Six divided by 20 is 0.300000, above the 0.10 pause threshold. Authorization and privacy still pass. Decision: pause. Preserve order and first attempts, assign an owner, and resume only through a new event after diagnosis. The incident card does not alter LAB-L06.

INC-04 occurs before a simulated restart. Authorization has expired. Projected restart cost is 12 units, exceeding the approved remaining eight units. Either is a hard-gate failure. Decision: stop and rollback to `SNAPSHOT-APPROVED-001`, the approved offline mode. No live request occurs, and unapproved rows are not admitted. Verify the panel hash and analyzer pass, then append rollback completion.

The event ledger records these decisions in sequence. The plan is row one. INC-01 is row two. INC-02 is row three. A planned B-02 look is row four. INC-03 is row five. INC-04 is row six. Rollback verification and closure are rows seven and eight. Hash-chain validation establishes that the frozen event list is internally consistent. It does not simulate the human and organizational complexity of a real incident.

Consider two counterfactuals. If INC-03 had produced one unavailable cell out of 20, the 0.05 value would not exceed the pause threshold; with hard gates and regime identity intact, the rule would continue. If INC-04 had sufficient budget but expired authorization, the decision would still stop and rollback because one hard-gate failure is enough. These counterfactuals show that precedence, not a narrative about incident severity, generates the result.

The rules also prevent outcome direction from entering incident decisions. A declining mention rate does not cause stop unless the plan defines an outcome guardrail. An improving rate does not prevent pause or stop. Monitoring remains tied to authorized design and evidence integrity rather than an optimization target.

## Chapter 10

We finish with evidence boundaries and the publication test.

PAPER-29 supports the importance of repeated observation in its particular Swiss-German, four-vertical, January–March 2026 study. Its run counts, window lengths, conditional metrics, and exclusions do not define a universal monitoring schedule. Same-window variation cannot be assigned solely to token sampling.

PAPER-38 is a single-author vendor preprint based on unreleased commercial data, differing provider policies, vendor-defined metrics, varying cohorts, and future protocol proposals. It does not provide independent reproduction or causal evidence for content interventions. Its black-box outputs do not disclose training, ingestion, memory, or retrieval causes.

PAPER-12 is a locked arXiv v2 fixed-slate testbed. PAPER-23 is a locked arXiv v2/KDD 2026 reconstructed pipeline. Both are valuable bounded environments; neither licenses production generalization. PLAT-04 is a named interface-documentation case, not a cross-engine metric standard or data-collection authorization.

LAB-L06 supports deterministic statements about 360 synthetic rows and its declared analysis. The incident simulator supports deterministic decisions under an authored rule. No named platform, hidden backend transition, causal treatment effect, detector performance, user behavior, or business outcome is estimated.

Use the final publication test. Did we record configuration, alias or explicit status, time, session, account, locale, order, failures, and releases? Did we separate scheduled outcome looks from continuous hard-gate monitoring? Does every change point say whether its cause is known or unknown? Were thresholds and precedence frozen before incident interpretation? Did authorization, privacy, rate limits, evidence preservation, policy, and cost override outcome convenience? Is the ledger append-only and hash-linked? Does rollback name and verify an approved target? Did we avoid inferring a hidden backend from the alias difference?

If any answer is no, the monitoring conclusion is not ready. If every answer is yes, the conclusion remains bounded: INC-01 continues, INC-02 rebaselines, INC-03 pauses, and INC-04 stops and rolls back under the supplied synthetic rules. That is procedural reproducibility, not a live-platform finding.

The final handoff package should allow a reviewer to reconstruct four layers separately: the frozen panel, the observable system-state cards, the append-only event history, and the decision procedure. If those layers are mixed, an incident card can be mistaken for observed panel data, or a release note can be mistaken for a causal explanation. Separation is a quality control.

Monitoring is not a promise that uncertainty will vanish. It is a disciplined way to preserve what was observed, identify what changed in the documented regime, expose what remains unobservable, and connect evidence to an authorized next action. Under change, that restraint is a form of learning.
