1. An observed change is not yet an intervention effect
Suppose citation incidence rises from 0.30 before a content edit to 0.42 afterward. The arithmetic change is 0.12. Calling it an effect requires a comparison world: what would the same eligible units have produced at the same time and system state if the intervention had not been applied? That outcome is not directly observed for treated units.
Several processes can produce the same before/after pattern: a platform release, index propagation, a different query mix, seasonality, concurrent source changes, competitor withdrawal, changing retry behavior, missing responses, or judge drift. A causal study must prevent, balance, measure, model, or explicitly assume away these alternatives. The method name alone does not do that.
W08 distinguishes four statements. A versioned observation reports what occurred under recorded conditions. An association compares observed groups or periods without a sufficient identification argument. An identified effect links a declared estimand to an assignment or quasi-experimental design plus defensible assumptions. A transported effect adds an argument that the result applies to a new population, system, intervention version, or time.
The appropriate response to weak identification is not to decorate a before/after difference with causal notation. Report the versioned association, list alternative explanations, and state the design needed to discriminate them.
2. Define the estimand before choosing an estimator
An estimand is the target contrast. A complete estimand card names:
- the treatment version and control version, including hashes;
- the unit whose potential outcomes are contrasted;
- the outcome, scale, eligibility, and observation window;
- the target population or finite set;
- the versioned system state or comparability window;
- the contrast, such as average assignment effect or effect among treated units;
- aggregation and weighting;
- treatment uptake and noncompliance policy;
- interference assumptions; and
- missing-outcome handling.
An estimator is the rule applied to observed data—for example, a difference in cluster means. An estimate is the resulting number. A correct computation can estimate the wrong target when units, weights, or eligibility do not match the estimand.
Let unit i have potential outcomes Y_i(1, S_t) and Y_i(0, S_t) under treatment and control at system state S_t. Only Y_i(D_i, S_t) is observed. A finite-set average effect is the mean of Y_i(1, S_t) - Y_i(0, S_t) across declared units. Potential-outcome notation makes the missing counterfactual visible; it does not supply it.
Treatment is the complete version. “Added a table” is inadequate if text, citations, markup, layout, and publication timing also changed. If several factors move, the estimand concerns the bundle unless a factorial, component, or sequential design separates them.
3. Use a DAG to expose paths and choices
A directed acyclic graph records causal assumptions. For a content intervention D, a stage outcome Y, platform state P, query mix Q, baseline content quality C, measurement rule M, and selection into observed responses R, plausible paths include C → D, C → Y, P → D, P → Y, Q → Y, D → R, and Y → R.
A variable that causes both treatment and outcome is a confounder. Blocking or randomization can break its treatment association; measured adjustment may close a backdoor path under assumptions. A mediator lies on the treatment pathway. Adjusting for it changes the target from a total effect toward a controlled or direct-effect question. A collider is caused by two variables; conditioning on it can open a spurious path. Complete-case selection can behave like collider conditioning when both treatment and outcome affect whether a response is observed.
Use different line styles for observed and latent variables. Mark time explicitly rather than treating “platform” as constant. Draw interference arrows between units when changing one page affects another source’s candidate capacity or competitive exposure.
A DAG does not prove that arrows are correct, that omitted nodes are irrelevant, or that adjustment suffices. It is a reviewable statement of assumptions. Pair it with measurement definitions, design evidence, temporal diagnostics, negative controls, and sensitivity analysis.
4. Choose the design family by assignment and failure mode
Individual randomization is attractive when independently assignable units receive versions without spillover. In GEO settings, queries, pages, domains, or time windows may be the real units. Repeated outputs are usually measurements, not assignment units.
Cluster randomization assigns intact groups such as query intents or domains. It protects against within-cluster contamination but needs enough independent clusters. Analysis must follow the cluster assignment.
Blocking or stratified randomization groups units by prognostic variables—intent, baseline rate, locale, or source type—then randomizes within blocks. It can improve balance and precision without changing the target if block weights are declared.
Matched designs pair treated and control units on pre-treatment variables when random assignment is unavailable. Matching does not balance unmeasured confounding. Match rules must be frozen before post-treatment outcomes and paired analyses must retain the match.
Staggered rollout assigns intervention timing across groups. It can create concurrent comparisons but introduces anticipation, carryover, dynamic effects, and heterogeneous timing. Modern estimators must match the rollout structure; a naive pooled before/after can mix incompatible contrasts.
Switchback or crossover alternates reversible versions by time or unit. It works only when carryover and propagation are negligible relative to the washout period and demand is comparable.
Difference-in-differences compares treated and untreated changes. It requires credible parallel trends absent treatment, no anticipatory response, stable measurement, and controlled spillover. Plot multiple pre-periods and placebo dates.
Interrupted time series estimates level or slope change after a dated intervention. It needs sufficient pre/post observations, stable measurement, a declared interruption, autocorrelation treatment, and evidence against concurrent shocks. A line break is not randomization.
5. Separate blocking, adjustment, and untestable assumptions
Blocking is a design action taken before outcomes: units are grouped and assignment occurs within those groups. Adjustment is an analysis action using measured covariates. Both can improve comparability, but neither is a generic cure.
Randomization supports exchangeability over its assignment mechanism when implemented and analyzed correctly. Blocking can increase precision and ensure balance on important variables. Covariate adjustment can improve precision or address measured confounding in an observational design. Yet all rely on treatment versions, uptake, valid outcomes, and appropriate unit handling.
Some assumptions remain only partly testable. Parallel pre-trends can make difference-in-differences more plausible but cannot prove post-period parallel trends. Balance checks can detect assignment failures but do not verify every latent variable. Negative controls can expose some pathways but cannot certify their absence. Sensitivity analysis asks how strong an unmeasured bias would need to be; it does not make the bias disappear.
Write an identification table with one row per threat: causal path, proposed design response, diagnostic, residual assumption, and verdict if the diagnostic fails. This separates evidence from hope.
6. Power, clustering, and repeated-measure dependence
Power planning starts with the estimand and assignment unit. Required inputs include the base rate or variance, minimum worthwhile effect, number of independent units, allocation, cluster sizes, intraclass correlation, repeated-run variance, time drift, attrition or missingness, multiplicity, and analysis model.
If each of thirty query intents is repeated five times, assignment by intent still yields thirty independent assignment units. The 150 rows estimate within-intent variability and can improve cluster means, but they do not create 150 independent treatment allocations. Treating every row as independent can produce falsely narrow intervals.
For roughly equal cluster size m and intraclass correlation rho, the diagnostic design effect is 1 + (m - 1)rho. It shows diminishing independent information from correlated repeats. It is not a universal sample-size correction. Unequal clusters, binary outcomes, few clusters, and complex random effects require simulation or design-specific methods.
Power is not permission to choose an implausibly large expected effect. Use a minimum decision-relevant effect and simulate the planned assignment, outcome distribution, missingness, and analysis. Power guardrails too: the study should detect severe fidelity or safety harm, not only a favorable primary outcome. When feasible sample size cannot support a useful decision, label the work a pilot, feasibility study, or descriptive audit.
7. Interference changes the potential-outcome object
Standard no-interference notation assumes unit i depends only on its own assignment. Competitive or networked information systems violate this easily. Editing source i can change source j’s rank, context slot, citation opportunity, or user attention. A domain-wide template can treat many pages simultaneously. Index propagation can contaminate controls. Reusable strategies can adapt based on earlier outcomes.
Potential outcomes then depend on an assignment vector or exposure mapping: Y_i(D_i, D_-i, S_t). The estimand must specify direct, spillover, total, displacement, or saturation effects. Cluster assignment may reduce cross-arm spillover when interference stays within clusters. Network or saturation designs may be needed when it does not.
Record candidate and context composition where the environment allows. Report target gain alongside displacement, source concentration, fidelity, and welfare guardrails. If contamination is discovered, preserve it in the deviation log and revise the identification verdict. Do not silently relabel contaminated controls.
8. Sequential peeking and multiplicity are design variables
Repeatedly checking outcomes and stopping at a favorable value inflates false-positive risk under conventional fixed-horizon inference. Freeze either a fixed sample/time horizon or a valid sequential rule with monitoring frequency, spending boundary, minimum information, and terminal decisions. Operational safety monitoring can continue, but distinguish harm stops from efficacy claims.
Multiple outcomes create another search space. Mention, citation, entailment, absorption, fidelity, referral, and risk are not interchangeable. Declare one primary outcome or a formal composite whose components and weights are justified. Mark secondary outcomes and guardrails. Control multiplicity for confirmatory claims or label the analysis exploratory.
Retries can become hidden sequential selection. Preserve the first attempt, reason, wait, count, and all terminal states. Never retry until a desired response appears. A refusal or missing citation panel may be an outcome, censoring state, or protocol failure depending on the estimand; decide before comparison.
9. Preregistration and deviation logs preserve the decision path
A useful preregistration freezes the research question, estimand, hypotheses, treatment hashes, assignment and blocking, unit, query frame, time window, outcomes, weights, missingness, retries, exclusions, sample-size rationale, primary estimator, uncertainty method, guardrails, stopping, multiplicity, sensitivity analyses, and release rule.
Preregistration does not guarantee a valid design. It makes undisclosed flexibility harder and provides a reference for deviations. Time-stamp it before outcome access and separate confirmatory from exploratory work.
A deviation log is append-only. Each entry records planned rule, observed departure, detection time, reason, affected units, whether outcomes were visible, corrective action, analysis consequence, approver, and status. Do not edit the original plan to make the deviation disappear. If platform state changes outside the comparability window, pause, rebaseline, or split the estimand rather than averaging across regimes.
10. Internal, construct, statistical, and external validity differ
Internal validity asks whether the observed contrast is attributable to the treatment for the studied units and period. Assignment integrity, uptake, confounding, interference, missingness, and measurement stability matter.
Construct validity asks whether treatment and outcome operationalizations represent the intended concepts. A citation counter is not automatically source influence; a layout change may bundle legibility and salience.
Statistical conclusion validity concerns precision, dependence, model fit, multiplicity, and error control. A small p-value cannot repair confounding, and wide intervals may make a sound design inconclusive.
External validity asks where the effect transports: systems, queries, locales, source types, intervention versions, dates, and institutional settings. A strong open reconstruction can have high internal transparency and limited closed-platform transport. A live field observation can have ecological realism and weak mechanism visibility.
Validity is a boundary statement, not a badge. A result can be excellent within a narrow target and honestly nonportable.
11. Read PAPER-10, PAPER-23, PAPER-32, and PLAT-04 by design
PAPER-10, CC-GSEO-Bench, proposes a content-centric benchmark and several influence dimensions over its query–article organization. W08 uses it to inspect units, aggregation, content-centric targets, and the distinction between exposure, credit, and causal impact. Reported benchmark results remain bounded to its construction and systems.
PAPER-23, SAGEO Arena, motivates end-to-end stage analysis in a reconstructed retrieval–reranking–generation environment. W08 asks how target selection, baseline filtering, pipeline reconstruction, and run uncertainty shape identification and transport. An open environment does not disclose a commercial platform.
PAPER-32, From Experience to Skill, is audit-only. Learners inspect benchmark size, test-set selection, strategy reuse, uncertainty, attribution claims, and possible adaptation or interference. Headline improvements are not admitted as course causal effects without claim-level verification of design and analysis.
PLAT-04 documents a named platform preview and its interface counts. It may support metric identity for that surface and date. It does not supply randomized assignment, counterfactual outcomes, source-use identification, or cross-platform transport.
12. Deterministic fixtures and the W08 identification verdict
L05 reproduces three locked claims, one changed factor, six equivalence checks, and zero software errors. Its ceiling is structural: no response outcome is observed. L06 reproduces 360 synthetic panel events, twenty queries, two teaching surfaces, three blocks, three repetitions, 354 complete events, and six noncomplete events. Its intervals are descriptive and do not adjust repeated-query dependence. No intervention is assigned.
Therefore L05 plus L06 does not equal a causal experiment. The package adds a separate eight-cluster pair-blocked teaching table. Four treated and four control clusters each have twenty binary opportunities in baseline and post periods. Treated average change is 0.1125; control average change is 0.0500; the difference-in-differences contrast is 0.0625. The arithmetic is reproducible. Its causal interpretation remains conditional on the frozen assignment, uptake, parallel counterfactual trend, no interference, stable measurement, and synthetic scope.
A defensible verdict has three parts: estimate, identification conditions, and transport boundary. If any central condition fails, report association or inconclusive evidence. The final question is not “Is the number positive?” It is “What comparison world, design, and assumptions make this number answer the declared estimand?”
13. Design-selection decision table and falsification checks
Choose a design by starting with authority and assignment feasibility, not by selecting the most sophisticated estimator. The following table is a compact review route:
| Situation | Preferred starting design | Evidence needed before effect language | Central falsification check |
|---|---|---|---|
| Many independent, authorized units can receive either version concurrently | randomized or block-randomized experiment | frozen versions, assignment log, uptake, cluster-aware analysis | balance and assignment-integrity audit |
| Units form natural spillover groups | cluster randomization | enough clusters, within-cluster exposure definition, spillover boundary | cross-cluster contamination check |
| Randomization is unavailable but strong pre-treatment matches exist | matched concurrent controls | match rule, overlap, pre-outcome covariates, residual-confounding analysis | hidden-bias sensitivity and pre-period outcome comparison |
| Rollout must occur, but timing can be assigned | randomized staggered rollout | timing assignment, cohort-specific effects, no anticipation, stable measurement | pre-rollout leads and cohort contamination |
| A reversible treatment can alternate quickly | switchback | washout, carryover model, stable time demand, balanced schedule | lagged treatment and period-order check |
| One treated series has a defensible concurrent comparison | difference-in-differences | several pre-periods, parallel-trend rationale, no differential shock | placebo dates and pre-trend contrasts |
| One series has a sharp intervention date but no control | interrupted time series | long pre-period, level/slope model, concurrent-event log | false interruption dates and residual autocorrelation |
| Only post-treatment convenience observations exist | descriptive association | precise sampling and measurement boundary | no causal verdict; redesign required |
Every effect claim should survive a treatment integrity audit. Confirm that assigned versions were actually served, no arm received an undeclared variant, propagation timing fits the window, and rollback or fallback did not silently change exposure. Intent-to-treat estimates preserve assignment even when uptake fails; per-protocol estimates answer a different, more selected question and require stronger assumptions.
Every outcome should survive a measurement equivalence audit. The parser, annotation guide, citation resolver, locale, and missingness logic must not change differentially by arm or time. Blind human adjudication where feasible. Use a negative-control outcome that should not respond to the intervention and a positive-control mechanism check that confirms the treatment can reach its hypothesized stage. A failed mechanism check can make a null effect uninterpretable because the intervention may not have been delivered.
Every temporal claim should survive a state-change audit. Record product releases, explicit model configurations where exposed, index or cache events, corpus changes, scheduled campaigns, collection incidents, and judge revisions. A change point does not identify its cause merely because it aligns with an intervention date. If a concurrent shock occurs, preserve the event and apply the preregistered pause, split, or rebaseline rule.
Finally, every decision should include an estimand-to-action map. Define what estimate supports retain, revise, rollback, or inconclusive; require fidelity, accessibility, safety, and subgroup guardrails; and prevent a favorable average from hiding severe harm. If the observed design cannot answer the decision-relevant estimand, the scientifically correct output is a redesign recommendation. More rows or a more complex model cannot substitute for the missing comparison.
14. Minimum analysis and reporting package
A confirmatory W08 report should permit reconstruction from assignment to verdict. Preserve a unit registry, block or match identifiers, version hashes, assignment seed or allocation record, exposure and uptake log, planned and observed cell flow, raw outcomes, missing-state table, exclusion and retry ledger, primary estimate with interval, guardrail table, subgroup results named in advance, and every deviation. The figure is secondary to this ledger.
Report the point estimate in outcome units. If the outcome is binary citation correctness, a difference of 0.06 is six percentage points, not “six percent” unless a relative contrast is explicitly intended. Show the arm-specific numerators and denominators, cluster counts, range of cluster sizes, and whether uncertainty follows clusters, matched pairs, or time blocks. With few clusters, avoid asymptotic confidence language that the design cannot support; display cluster-level contributions and randomization-compatible reasoning where possible.
Also report treatment delivery as a flow rather than a footnote. Count units assessed for eligibility, assigned, exposed to the intended version, measured, excluded, and analyzed. Preserve crossovers and failed uptake in their assigned arms for the intent-to-treat primary estimate, then label any uptake-based analysis as a different estimand. A flow table makes denominator changes visible and prevents silent removal of difficult clusters.
Include at least three sensitivity routes: one for missing outcomes, one for analysis specification, and one for comparability or spillover. A missingness sensitivity may count every noncomplete treated event as failure and every noncomplete control event as success, then reverse the assignment, to bound a binary contrast. A specification sensitivity may compare prespecified unadjusted and covariate-adjusted estimates. A spillover sensitivity may remove contaminated clusters according to a frozen rule while retaining the intent-to-treat primary analysis.
End with a validity paragraph written as claims, not boilerplate: which units were exchangeable by design; which paths remained assumed; whether treatment uptake and measurement were verified; what interference was possible; what the interval excludes; and where transport is unsupported. A transparent null or inconclusive result that preserves this chain has more scientific value than a positive contrast whose comparison world cannot be reconstructed. Identification quality begins with reconstructible decisions and evidence.