# W08 Slide Script — From Observation to Identification

Production note: exactly twenty-four slides are specified. All visuals are original instructional diagrams. Color may reinforce grouping, but labels, geometry, pattern, and line style must carry every distinction.

## Slide 01 — After is not because

**On-screen text**

Before: 0.30. After: 0.42. Observed change: +0.12.

Missing: what would have happened after without treatment?

**Visual specification**

Two observed bars sit above an empty outlined bar labelled `post-period under control`. Seven dashed arrows from platform state, query mix, seasonality, concurrent edits, propagation, missingness, and judge drift point toward the observed change.

**Speaker notes**

Begin with the arithmetic and refuse causal language. The missing counterfactual is not the baseline bar. Ask which alternative processes could produce the rise. W08 is about designing a comparison world, not making a positive number look sophisticated. A causal verdict needs treatment versions, assignment, uptake, outcome, estimand, estimator, and assumptions.

**Teaching check**

Learners name three explanations other than treatment and rewrite “the edit improved citation” as a versioned association.

**Alt text**

“Observed before and after bars surround a missing post-control counterfactual while seven alternative causes point to the change.”

## Slide 02 — The estimand card

**On-screen text**

Treatment · control · unit · outcome · population · system/time · contrast · weights · uptake · interference · missingness

**Visual specification**

An eleven-row card places treatment and control versions at the top, the unit and target in the middle, and assumptions at the bottom. Hash tags appear beside both versions. A red empty-field indicator blocks the estimate arrow.

**Speaker notes**

An estimand is the target contrast. It must exist before choosing a statistical procedure. Treatment means the whole version, not a slogan such as “added evidence.” The card should identify what is averaged and where it applies. If a central field is empty, the study does not yet know what effect it seeks.

**Teaching check**

Give “Does better structure work?” and require a complete unit, outcome, population, and time-specific contrast.

**Alt text**

“An eleven-field estimand card blocks analysis until treatment, control, target, uptake, interference, and missingness are specified.”

## Slide 03 — Estimand, estimator, estimate

**On-screen text**

Estimand: target. Estimator: rule. Estimate: realized number.

Correct arithmetic can answer the wrong question.

**Visual specification**

Three differently shaped objects form a pipeline: a target reticle, a calculation machine, and a numerical receipt. Two mismatch arrows show the machine receiving the wrong unit weights and producing a precise but off-target receipt.

**Speaker notes**

Use a cluster-average effect as the target, a difference in cluster means as the estimator, and 0.0625 as one estimate. If rows rather than clusters are weighted, the estimator may target a response-weighted contrast. Precision does not fix target mismatch. Every reported estimate must point backward to an estimand card.

**Teaching check**

Learners classify “average assignment effect,” “difference in means,” and “six-point-two-five percentage points.”

**Alt text**

“Target, calculation rule, and numerical result are separate objects, with mismatched weighting producing a precise off-target estimate.”

## Slide 04 — Potential outcomes expose the missing cell

**On-screen text**

`Y_i(1, S_t)` and `Y_i(0, S_t)`

Observe one per unit. System state belongs in the object.

**Visual specification**

Each of four unit rows contains two potential-outcome cells; one is solid observed and the other is hatched missing. A vertical system-state strip `S_t` touches every row. No imputation arrow fills the missing cells.

**Speaker notes**

Potential-outcome notation states the comparison but does not reveal individual counterfactuals. The versioned system state matters because platform or corpus changes can alter both potential outcomes. Identification comes from assignment or a quasi-experimental argument, not from the notation itself.

**Teaching check**

Why can one run not reveal both potential outcomes for the same unit? What would transport from `S_t` to `S_t+1` require?

**Alt text**

“Four units each reveal one of two treatment-state outcomes while the alternative remains hatched under a shared versioned system state.”

## Slide 05 — A DAG is an assumption map

**On-screen text**

Content quality → treatment choice  
Content quality → outcome  
Treatment → outcome  
Platform state and query mix → outcome

DAG ≠ proof.

**Visual specification**

A directed graph uses solid boxes for measured nodes and hatched boxes for latent nodes. Backdoor paths are bracketed; the treatment-to-outcome arrow is bold but not labelled causal until a design badge is attached.

**Speaker notes**

The graph makes proposed paths reviewable. Baseline content quality can confound an observational comparison if stronger pages are more likely to receive treatment and also perform better. Platform state and query mix require time and sampling control. A DAG cannot verify omitted-variable absence or arrow direction.

**Teaching check**

Learners identify the backdoor path through content quality and propose randomization, blocking, or measured adjustment.

**Alt text**

“A measured-and-latent DAG displays a content-quality backdoor path and labels the graph as an assumption map rather than proof.”

## Slide 06 — Confounder, mediator, collider

**On-screen text**

Confounder: common cause. Mediator: treatment pathway. Collider: common effect.

Adjustment changes paths—and sometimes the estimand.

**Visual specification**

Three mini-DAGs appear side by side with distinct arrow patterns. A gate icon closes the confounder path after appropriate control; scissors on the mediator indicate a direct-effect target; opening doors at the collider warn against conditioning.

**Speaker notes**

Do not build a covariate list by predictive strength alone. Pre-treatment confounders may need design or adjustment. Mediators should remain for a total-effect estimand. Conditioning on a collider can create association. Response availability can be a collider when treatment and outcome both affect capture.

**Teaching check**

Ask why complete-case analysis can create selection bias when both treatment and response quality affect observability.

**Alt text**

“Three small graphs contrast closing a confounder path, cutting a mediator pathway, and accidentally opening a collider path.”

## Slide 07 — Randomize at the real unit

**On-screen text**

Assignment unit may be query intent · page · domain · time block.

Repeated answers are nested measurements.

**Visual specification**

Thirty large query-cluster cards split into treatment and control. Each card contains five small response dots. A denominator label reads `30 assignment units, 150 measurements`, with arrows staying at cluster level.

**Speaker notes**

Randomization creates an assignment mechanism, not independent rows. If treatment is assigned by query intent, five repeated calls remain within that intent. Uncertainty must respect assignment. Randomizing individual prompts inside a semantic cluster may contaminate closely related versions.

**Teaching check**

Thirty intents each have five repeats. How many independently assigned units exist? Why does a sixth repeat not replace ten lost intents?

**Alt text**

“Thirty cluster cards contain 150 response dots, emphasizing thirty independent assignments rather than 150 independent units.”

## Slide 08 — Blocking improves balance before outcomes

**On-screen text**

Block on prognostic pre-treatment variables. Randomize within blocks.

Declare block weights.

**Visual specification**

Four intent strata form horizontal lanes. Each lane contains paired baseline cards and a balanced treatment/control split. A separate analysis scale combines block estimates with printed weights.

**Speaker notes**

Blocking is a design action. It ensures balance on important observed characteristics and can improve precision. It does not solve unmeasured confounding outside a randomized design. Blocks and weights must be defined before outcomes. If one block has severe harm, a favorable pooled average should not hide it.

**Teaching check**

Choose two plausible blocking variables for a multilingual query study and state why post-treatment citation count is not one.

**Alt text**

“Four pre-treatment strata each randomize treatment and control, then combine block estimates through declared weights.”

## Slide 09 — Matching is not randomization

**On-screen text**

Match on pre-treatment overlap. Preserve pairs. Audit residual bias.

Unmeasured confounding remains.

**Visual specification**

Treated and untreated unit cards align in pairs by baseline score, intent, and locale. Several unmatched cards fall outside a common-support frame. A dotted latent-confounder arrow bypasses the matching variables.

**Speaker notes**

Matching can construct a more comparable observational sample, but it changes the target toward units with overlap and cannot balance unmeasured causes. Freeze the distance, caliper, replacement rule, and tie policy before post-treatment outcomes. Analyze matched structure rather than discarding it.

**Teaching check**

What changes when high-baseline treated units have no comparable controls? Learners name overlap loss and target-population narrowing.

**Alt text**

“Pre-treatment matched pairs sit inside common support while unmatched units and a latent confounder remain visibly outside.”

## Slide 10 — Staggered rollout needs timing logic

**On-screen text**

Assign adoption time. Check anticipation · cohort effects · carryover · changing state.

Never-treated and not-yet-treated are not always interchangeable.

**Visual specification**

A unit-by-time grid shows four cohorts adopting treatment at different columns. Pre-treatment cells are patterned, post-treatment cells solid, and anticipation windows outlined. A platform-state band changes midway across all cohorts.

**Speaker notes**

Staggering can create concurrent comparisons when immediate universal rollout is unnecessary. Dynamic effects and heterogeneous cohorts complicate pooled estimators. A platform release can coincide with later cohorts. Define adoption, wash-in, control eligibility, and cohort-specific contrasts.

**Teaching check**

Why can an already-treated cohort be a poor control for a newly treated cohort? It may carry treatment effects and face a different state.

**Alt text**

“A staggered adoption grid marks cohort treatment timing, anticipation windows, and a platform-state change crossing every cohort.”

## Slide 11 — Difference-in-differences is an assumption, not subtraction alone

**On-screen text**

`(treated post − treated pre) − (control post − control pre)`

Requires credible untreated parallel trend.

**Visual specification**

Two pre-period trend lines run in parallel, then diverge after intervention. A second small panel shows nonparallel pre-trends and a crossed-out DiD badge. Placebo dates appear as vertical dotted lines.

**Speaker notes**

The formula removes a common observed change only if controls represent the treated counterfactual trend. Multiple pre-periods support diagnostics, not proof. Check anticipation, spillover, stable outcome measurement, composition, and concurrent shocks. Use placebo dates and alternative controls.

**Teaching check**

If treated and control pre-trends diverge before intervention, what verdict is justified? Redesign or an association with explicit limitation.

**Alt text**

“Parallel pre-trends diverge after treatment in one panel, while a nonparallel counterexample rejects the same subtraction as identification.”

## Slide 12 — Interrupted time series needs a credible interruption

**On-screen text**

Estimate level and slope change. Model autocorrelation. Log concurrent events.

A line break is not assignment.

**Visual specification**

A long time series has twelve pre and eight post points, fitted pre-trend extension, level shift, and slope change. Three concurrent-event flags sit at the intervention date; a falsification panel tests two fake dates.

**Speaker notes**

ITS is useful when one intervention date exists and no comparable control is feasible. It needs sufficient history, stable measurement, a justified functional form, and residual diagnostics. A simultaneous product update or campaign defeats clean attribution. Segmented regression is an estimator, not the identification argument.

**Teaching check**

Name two concurrent events that should enter the incident log and one placebo-date diagnostic.

**Alt text**

“A segmented time series separates projected trend, level, and slope while concurrent shocks and fake interruption dates test the design.”

## Slide 13 — Interference changes the outcome function

**On-screen text**

`Y_i(D_i, D_-i, S_t)`

Direct gain · spillover · displacement · saturation.

**Visual specification**

Four source nodes compete for three context slots. Treating source A moves source B out; arrows connect each unit’s assignment to neighboring outcomes. A cluster boundary encloses two strongly connected sources.

**Speaker notes**

When finite retrieval or context capacity creates competition, one unit’s treatment can alter another unit’s exposure. A template edit may treat many pages. Define an exposure mapping or cluster assignment. Report displacement and guardrails, not target gain alone. Contaminated controls remain in the deviation record.

**Teaching check**

What estimand changes when treatment to A removes B from context? Learners distinguish A’s direct effect from system-level displacement.

**Alt text**

“Four sources compete for three slots, with treatment of one displacing another and creating cross-unit outcome arrows.”

## Slide 14 — Power, clustering, and ICC

**On-screen text**

`DE = 1 + (m − 1)ρ`

Plan: base rate · minimum worthwhile effect · clusters · repeats · missingness · guardrails.

**Visual specification**

A staircase shows information gained from new clusters rising faster than information from repeated dots inside the same cluster. Three curves for low, medium, and high intraclass correlation diverge as repeats increase.

**Speaker notes**

The design-effect formula is a diagnostic under roughly equal clusters. Positive within-cluster correlation reduces the independent information in extra repeats. Simulate the actual assignment and binary outcome when planning. Power a meaningful decision, not the effect size needed to make a convenient sample pass.

**Teaching check**

If `m=5` and `rho=0.25`, compute design effect 2.0 and explain why it is not a universal correction.

**Alt text**

“New clusters add more independent information than correlated repeats, with higher intraclass correlation steepening the design-effect curves.”

## Slide 15 — Missingness belongs in the estimand

**On-screen text**

Refused · missing · parser failure · unavailable panel · policy block

Outcome, censoring, retry, or protocol failure—decide first.

**Visual specification**

A planned-cell flow separates complete, refused, missing, and invalid branches. No branch disappears. Two worst-case sensitivity trays invert missing treated/control outcomes. A retry ledger preserves first attempt.

**Speaker notes**

Complete-case analysis changes the target when observability depends on treatment or outcome. Classify states before comparison. Bound retries and retain every attempt. Show response rates by arm and time. Use sensitivity bounds when binary outcomes and small missing counts permit transparent worst cases.

**Teaching check**

Why is a missing response not automatically zero citation? It may be a different event or censoring process under the protocol.

**Alt text**

“Every planned cell flows to an explicit terminal state, with worst-case missingness bounds and a preserved retry ledger.”

## Slide 16 — Sequential peeking changes error rates

**On-screen text**

Fixed horizon—or declared sequential rule.

Monitoring cadence · boundary · minimum information · stop actions.

**Visual specification**

Ten cumulative estimate points cross an ordinary threshold on the fourth look and later return. A second line uses a wider sequential boundary. Separate red safety stops remain active throughout.

**Speaker notes**

Repeatedly looking and stopping at the first favorable value invalidates ordinary fixed-horizon claims. Predeclare a fixed sample or a valid sequential design. Safety monitoring is different from efficacy stopping and may require immediate action. Archive every look and decision.

**Teaching check**

Correct “we stopped when p fell below .05” by naming what would have needed to be declared.

**Alt text**

“A cumulative estimate briefly crosses a fixed threshold but not a sequential boundary, while independent safety stops remain available.”

## Slide 17 — One primary outcome, explicit guardrails

**On-screen text**

Primary: decision target. Secondary: explanation. Guardrails: must not worsen.

Multiplicity is not a menu for a positive result.

**Visual specification**

One large primary-outcome card sits above four secondary cards and three red guardrail gates for fidelity, accessibility, and risk. A locked decision rule connects them to retain, revise, rollback, or inconclusive.

**Speaker notes**

Mention, citation, entailment, absorption, referral, and safety measure different stages. Declare the primary or formal composite, secondary analyses, and guardrails. A favorable mean cannot override a severe guardrail breach. Control multiplicity for confirmatory claims or label exploratory findings clearly.

**Teaching check**

Learners choose a primary outcome for a citation-support intervention and two non-negotiable guardrails.

**Alt text**

“One primary outcome, four secondary outcomes, and three mandatory guardrail gates feed a frozen four-way decision rule.”

## Slide 18 — Preregistration freezes choices, not truth

**On-screen text**

Freeze: estimand · versions · assignment · sample · outcomes · missingness · analysis · stops · release rule.

Time-stamp before outcome access.

**Visual specification**

A preanalysis document receives a timestamp seal. Nine tabs show required sections. A transparent lock covers the plan while a separate exploratory notebook remains open but distinctly labelled.

**Speaker notes**

Preregistration limits undisclosed flexibility. It does not make an invalid design valid. Freeze hashes and assignment along with hypotheses. Separate confirmatory and exploratory analyses. Reviewers should reconstruct when outcomes became visible and which decisions preceded them.

**Teaching check**

Name three choices that cannot be safely finalized after viewing arm differences.

**Alt text**

“A timestamped nine-tab preanalysis plan is locked before outcomes, while a separate notebook preserves labelled exploration.”

## Slide 19 — The deviation log is append-only evidence

**On-screen text**

Planned rule · departure · detection · outcomes visible? · affected units · action · analysis impact · approver

Do not rewrite history.

**Visual specification**

An immutable event ledger shows three deviations: propagation delay, parser change, and control contamination. Each branches to continue, pause, rebaseline, or invalidate. The original plan remains visible beneath the ledger.

**Speaker notes**

Operational studies rarely follow every plan perfectly. Scientific integrity comes from preserving departures and their consequences. If a product state changes beyond the comparability window, the correct action may be pause or split, not smoothing. Never edit the preregistration to make a late decision appear planned.

**Teaching check**

Write one deviation entry for an outcome parser update discovered after half the sample.

**Alt text**

“An append-only deviation ledger keeps the original plan visible and routes each departure to a documented decision.”

## Slide 20 — Four validity boundaries

**On-screen text**

Internal · construct · statistical conclusion · external

Strong in one dimension does not imply strong in all.

**Visual specification**

Four nested but nonidentical frames surround a study: assignment, measurement construct, uncertainty, and transport population. Gaps between frames are labelled with remaining threats. No single quality badge appears.

**Speaker notes**

Internal validity concerns attribution within the studied setting. Construct validity asks whether “citation” or “influence” was measured as intended. Statistical conclusion validity concerns precision and error. External validity concerns transport. An open reconstruction may be internally transparent yet poorly transport to a closed surface.

**Teaching check**

Classify “the dashboard counter does not measure absorption” and “the result may not apply next year” into validity dimensions.

**Alt text**

“Four separate study frames mark assignment, construct, uncertainty, and transport boundaries without collapsing them into one badge.”

## Slide 21 — L05 validates a treatment package, not an effect

**On-screen text**

3 locked claims · 3 source identities · 1 changed factor · 6 equivalence checks · 0 errors

No response outcome measured.

**Visual specification**

Control and treatment pages align claim by claim. Only a stylesheet factor card differs. Hash and rollback cards sit below. A closed outcome chamber at right is labelled `not collected`.

**Speaker notes**

L05 demonstrates version control, factual equivalence, authorization, and reversibility in a local fixture. Its declared factor is evidence-layout proximity and hypothesized stage is human comprehension. A passing audit cannot establish crawling, ranking, citation, or any causal visibility outcome.

**Teaching check**

State the strongest L05 result and one effect sentence it cannot support.

**Alt text**

“Hash-aligned control and treatment pages differ by one declared style factor while the outcome chamber remains explicitly unobserved.”

## Slide 22 — L06 describes a panel, not treatment assignment

**On-screen text**

360 events · 20 queries · 2 synthetic surfaces · 3 blocks · 3 repeats  
354 complete · 6 noncomplete

Descriptive Wilson intervals; repeated-query dependence remains.

**Visual specification**

A four-dimensional cube maps query, surface, block, and repetition. Six hollow cells mark noncomplete events. A descriptive interval chart is fenced away from an empty treatment-assignment card.

**Speaker notes**

L06 preserves a complete designed panel and separate outcomes. It exposes missingness and drift. It contains no assigned content intervention, so surface or time contrasts are not intervention effects. Its Wilson intervals do not adjust clustered repeated observations. Combining L05 and L06 informally would not fill the missing assignment link.

**Teaching check**

Why does “L05 treatment plus L06 change” fail identification? They are not one joined predeclared design.

**Alt text**

“A complete synthetic panel cube contains six missing cells and descriptive intervals, separated from a blank intervention-assignment record.”

## Slide 23 — Reproduce the W08 synthetic contrast

**On-screen text**

4 treated clusters: mean change `0.1125`  
4 controls: mean change `0.0500`  
Difference-in-differences: `0.0625`

Arithmetic reproduced; causal interpretation conditional.

**Visual specification**

Four matched pairs each show twenty baseline and post opportunities. Change arrows feed two mean-change trays and a subtraction ledger. Six assumption gates—assignment, uptake, parallel trend, no interference, stable measurement, synthetic scope—follow the number.

**Speaker notes**

Calculate cluster rates, then average cluster changes equally. The result is six-point-two-five percentage points. It is a deterministic teaching contrast, not a live engine effect. The gates state what would be required for an identification claim even within the synthetic design.

**Teaching check**

Learners reproduce `0.1125 − 0.0500 = 0.0625` and name the assumption most vulnerable to competitive spillover.

**Alt text**

“Eight cluster changes yield a 0.0625 difference-in-differences contrast followed by six unpassed identification gates.”

## Slide 24 — Exit: issue an identification verdict

**On-screen text**

Estimate · design evidence · assumptions · diagnostics · transport boundary

Verdict: identified within scope · association · inconclusive · invalid.

Source boundary: `PAPER-10` / `PAPER-23` core · `PAPER-32` audit-only · `PLAT-04` interface-only.

**Visual specification**

A five-field verdict card feeds four labelled exits. Behind it, the opening missing counterfactual is now connected to an assignment-and-control bridge, but the external transport landscape remains beyond a dashed boundary.

**Speaker notes**

Close with a complete sentence: report the contrast, units, design, assumptions, and scope. If assignment or comparability fails, downgrade the claim. A null result can be informative when treatment uptake is verified and intervals exclude worthwhile effects; otherwise it may be inconclusive. The goal is a reconstructible decision, not mandatory causality.

**Teaching check**

Final prompt: turn one before/after claim into an estimand and choose the strongest justified verdict.

**Alt text**

“A five-field evidence card routes the result to identified, associated, inconclusive, or invalid, while transport remains outside the study boundary.”
