# W07 Slide Script — Metrics, Uncertainty, and Repeated Observation

## Production conventions

- **Format:** 16:9 academic deck in deep purple, warm ivory, graphite, and restrained gold. Color is never the only state carrier.
- **Typography:** at least 30 pt body and 44 pt title. Numeric tables use aligned digits and voiced endpoints.
- **Visual authorship:** all visuals are original instructional diagrams. Do not reuse a paper figure, product screenshot, PLAT-04 interface capture, repository image, or platform logo.
- **Evidence grammar:** solid border means observed fixture data; double border means frozen design; dashed line means sensitivity; hatched area means unknown or not evaluated.
- **Accessible production:** preserve the exact alt text, retain table equivalents, label every interval endpoint, and avoid animation-dependent meaning.

## Slide 01 — When is a change repeatable?

**On-screen text**

> W07 · Metrics, uncertainty, and repeated observation  
> Essential question: When is a visibility change repeatable?

Footer: “One result is an event; repeatability is a design.”

**Visual specification**

One isolated response card appears at left. At right, a nested panel contains query rows, two surface columns, three time blocks, and three repeat dots per cell. A double border marks the frozen design; no commercial interface is depicted.

**Speaker notes**

Begin by contrasting a screenshot with a panel. One response can document one visible state. Repeatability requires a metric contract, repeated cells, time structure, and missing-state rules. State that this week studies an offline synthetic fixture. It neither queries nor represents a named live engine.

**Teaching check**

Learners name three design fields missing from the sentence “visibility increased.”

**Alt text**

“A single response card is contrasted with a nested query-by-surface-by-time panel containing repeated observations.”

## Slide 02 — Outcomes and audit artifacts

**On-screen text**

By the end, you can produce:

1. a complete metric card;
2. a validated repeated panel;
3. separate run-variation and drift summaries;
4. a query-cluster interval;
5. a predefined sensitivity table;
6. a bounded repeatability conclusion.

**Visual specification**

Six numbered output rows align with small labeled sheets: metric card, panel key, two-scale table, bootstrap manifest, sensitivity grid, and boundary memo. A final bracket reads “reconstructable without oral explanation.”

**Speaker notes**

Emphasize evidence artifacts over vocabulary recall. Every estimate should resolve to raw designed cells and a versioned metric card. The bootstrap must include cluster, seed, replicates, statistic, and endpoint rule. The conclusion must keep synthetic reproduction separate from live empirical claims.

**Teaching check**

Ask which artifact records whether missing responses were excluded or scored in a compound outcome.

**Alt text**

“Six learning outcomes align with six audit artifacts and a reconstruction requirement.”

## Slide 03 — The metric card

**On-screen text**

Metric card fields:

`event · numerator · denominator · unit · query frame · weights · surfaces · blocks · repeats · retry · missing states · aggregation · uncertainty · ceiling`

Write before estimating.

**Visual specification**

A large index card uses fourteen labeled rows. The numerator and denominator rows are joined by a fraction line; cluster and uncertainty rows are linked. A version badge and freeze icon use text labels as well as shape.

**Speaker notes**

A display name is not a protocol. Changing the denominator, weight, retry rule, or aggregation changes the metric version. Freezing the card before comparative outputs blocks convenient post-result redefinition. Invite learners to find the field most often omitted in dashboards: the eligible denominator and its missingness.

**Teaching check**

Learners add the missing fields to “citation rate equals cited answers divided by all answers.”

**Alt text**

“A fourteen-field metric card links numerator, denominator, cluster, uncertainty, and claim ceiling before estimation.”

## Slide 04 — Six events, six units

**On-screen text**

| Event | Example unit |
|---|---|
| Mention | entity–response |
| Citation | source–response |
| Entailment | claim–passage–response |
| Absorption | distinctive claim–response |
| Referral | user–source event |
| Action | user–task outcome |

No silent substitution.

**Visual specification**

Six ladder rungs carry event names and units. Conditional arrows connect rungs, while gaps show that one event does not imply the next. Each rung has a different written unit, not merely a different color.

**Speaker notes**

Explain how conditioning changes denominators. Entailment among resolved fetched citations is not citation incidence among complete responses. Referral may require exposure eligibility. A composite across these events lacks a natural unit unless a declared decision problem supplies weights and tradeoffs.

**Teaching check**

Ask whether citation frequency can serve as absorption without a source-distinctive claim set and answer-use rule.

**Alt text**

“Six stage-aligned event rungs use different analysis units and conditional gaps to prevent metric substitution.”

## Slide 05 — Denominator flow

**On-screen text**

Designed cells → returned responses → citation-bearing responses → resolved citations → fetched passages → adjudicated claim pairs

Every rate names its conditioning set.

**Visual specification**

A six-stage narrowing flow prints a count placeholder at every boundary. A side arrow warns that skipping from adjudicated pairs back to designed cells changes the estimand. Missing rows branch into labeled status boxes instead of disappearing.

**Speaker notes**

The denominator is not whichever count is convenient. A page-quality score among successfully fetched citations targets a fetch-conditioned population. Replacing that denominator with all prompts makes a different claim. Preserve the flow so an analyst can locate exclusions, failures, and undefined outcomes.

**Teaching check**

Learners identify the denominator for citation correctness versus citation incidence.

**Alt text**

“A denominator funnel narrows from designed cells to adjudicated claim pairs while missing states branch visibly at each stage.”

## Slide 06 — Build the Cartesian panel

**On-screen text**

L06 design:

`20 queries × 2 surfaces × 3 blocks × 3 repetitions = 360 cells`

Key: `query_id, surface, time_block, repetition`  
Every designed combination remains a row.

**Visual specification**

A small three-dimensional matrix is flattened into a table. Query bands contain surface columns; each surface contains three block cells; each block contains three numbered repeat dots. The equation and 360 total appear below.

**Speaker notes**

Validate the design before outcomes. Unique combination keys distinguish a complete panel from a file that silently lost rows. L06 preserves noncomplete cells. The two surfaces are named SYN-A and SYN-B to block accidental product attribution. It never estimates an external population.

**Teaching check**

Ask how many designed rows remain if a response is missing. The answer remains 360; status changes, not design count.

**Alt text**

“Twenty query bands nest two surfaces, three time blocks, and three repetitions for a total of 360 designed cells.”

## Slide 07 — Query strata and weights

**On-screen text**

Fixture strata: compare · learn · select · verify  
Five queries each.

Equal diagnostic weight ≠ target demand weight.

`estimate = Σ target_weight_h × stratum_mean_h`

**Visual specification**

Four equal fixture boxes sit above a separate target-weight bar with unequal labeled segments. An arrow reads “new estimand, not correction.” Both include numeric weights rather than color-only proportions.

**Speaker notes**

A balanced panel supports comparable diagnostics across intent strata. It does not estimate real query prevalence. Target weights may come from a sampling design, decision population, or independent usage data. Declare their source before seeing the contrast. Never call priority weights population demand.

**Teaching check**

Learners explain why equal counts per stratum do not justify twenty-five percent target weights.

**Alt text**

“Four equally sized fixture strata are contrasted with an unequally weighted target distribution labeled as a different estimand.”

## Slide 08 — Aggregation order changes weights

**On-screen text**

Possible estimators:

- complete-row mean
- equal-query mean
- equal-stratum mean
- target-weighted stratum mean
- designed-event compound rate

Same panel; different questions.

**Visual specification**

One event table feeds five calculation paths. Each path groups rows in a different sequence and ends in a distinctly labeled estimate box. No path is marked inherently preferred.

**Speaker notes**

If missingness differs by query, the complete-row mean gives more weight to queries with more observed repetitions. Equal-query analysis first creates one query summary. A designed-event compound rate scores noncompletion into a response-and-event outcome. Each needs its own name and target.

**Teaching check**

Ask when a row mean and equal-query mean coincide and why coincidence does not erase their structural difference.

**Alt text**

“Five aggregation paths transform the same event table into differently weighted, explicitly named estimands.”

## Slide 09 — Two time scales

**On-screen text**

**Within block:** repeated-run disagreement under nominally similar conditions.  
**Between blocks:** change across authored time states.

Do not pool before inspecting both.

**Visual specification**

Three large block columns each contain three close repeat dots for one query. Small brackets measure within-block spread. A longer arrow connects first and last block means, labeled descriptive drift.

**Speaker notes**

Repeated runs estimate run-level variation inside a cell. Later blocks observe a different time state. A one-hour repeat schedule cannot establish month-scale stability. A block difference cannot identify whether a platform, corpus, collection, or authored fixture mechanism changed. Keep their records separate.

**Teaching check**

Learners classify ten same-hour runs and three monthly observations by the variation component they can inform.

**Alt text**

“Close repeat dots within each of three time blocks are separated from a longer first-to-last drift arrow.”

## Slide 10 — One unstable cell

**On-screen text**

Q-001 · SYN-A · B-01 · mention

Repeats: `[1, 1, 0]`  
Mean: `2/3`  
Sample variance: `1/3`

One query cluster; three measurements.

**Visual specification**

Three numbered repeat circles display one, one, and zero. They sit inside one double-bordered query cell. A bracket labels the cell mean and sample variance; an outer label states cluster count one.

**Speaker notes**

The disagreement shows why one run is a fragile description. Three observations do not create three sampled queries. The sample variance is a within-cell statistic, not a between-block drift measure or population interval. Preserve the raw sequence rather than only its mean.

**Teaching check**

Ask how many query clusters and how many repeat measurements appear on this slide.

**Alt text**

“One query cell contains three binary repeats, one, one, and zero, with mean two thirds and sample variance one third.”

## Slide 11 — First-to-last block drift

**On-screen text**

SYN-A mention:

`B-01: 28/58 = 0.482759`  
`B-03: 24/60 = 0.400000`  
`last − first = −0.082759`

Descriptive fixture change; mechanism unknown.

**Visual specification**

Three block points form a simple line with printed numerators, denominators, and rates. A hatched band behind the line reads “cause not identified.” The intermediate B-02 point is 0.423729.

**Speaker notes**

This exact result comes from the frozen panel. The changing denominator reflects missing statuses. The last-minus-first arithmetic is reproducible, but it cannot attribute the change to a release, corpus update, session state, or any live process. Three authored blocks are not a general time-series study.

**Teaching check**

Learners state the observed change and the strongest mechanism claim that remains unsupported.

**Alt text**

“SYN-A mention declines from 28 of 58 to 24 of 60 across three blocks, with the cause shown as unknown.”

## Slide 12 — Dependence lives inside the panel

**On-screen text**

Rows share:

- query wording and intent
- surface state
- time block
- collection configuration

`360 rows ≠ 360 independent query intents`

Primary cluster for equal-query inference: 20 query IDs.

**Visual specification**

The 360 event dots are gathered into 20 labeled query envelopes. Within each envelope, lines connect repeated surface/block events. A counter at bottom reads twenty clusters and 360 measurements.

**Speaker notes**

Repeated rows add information about within-query behavior but do not expand the query frame. Treating every row as independent can narrow intervals unjustifiably. The correct resampling or model unit follows the sampling or assignment design and the target statistic. That distinction governs uncertainty reporting.

**Teaching check**

Ask whether a fourth repetition can compensate for losing five independently sampled query intents.

**Alt text**

“Three hundred sixty event dots are nested inside twenty query envelopes, distinguishing measurements from primary query clusters.”

## Slide 13 — Wilson intervals: useful and bounded

**On-screen text**

Example: SYN-A · B-01 · mention

`28/58 = 0.482759`  
Wilson 95%: `[0.359282, 0.608377]`

L06 label: descriptive; not dependence-adjusted.

**Visual specification**

A horizontal interval line has numeric endpoints and a center marker at the observed rate. Beneath it, four omitted layers are written: query-frame uncertainty, dependence, label error, and drift mechanism.

**Speaker notes**

Wilson intervals describe a binary proportion under the supplied denominator more safely than some simple normal approximations. They do not solve the panel’s dependence. Interval overlap is not a formal equality test, and a narrow interval does not repair a biased query frame.

**Teaching check**

Learners name two sources of uncertainty not represented by the displayed Wilson endpoints.

**Alt text**

“A Wilson interval from 0.359282 to 0.608377 surrounds rate 0.482759 while four omitted uncertainty sources remain listed.”

## Slide 14 — Query-cluster bootstrap algorithm

**On-screen text**

1. Calculate one SYN-B minus SYN-A mention difference per query.
2. Sample 20 query IDs with replacement.
3. Retain each selected query statistic.
4. Average; repeat 5,000 times.
5. Sort and apply declared endpoint indices.

Seed: 707.

**Visual specification**

Twenty query cards feed a seeded sampling wheel. One example draw includes a duplicated query card and an omitted card. The wheel outputs one mean; 5,000 means form a sorted strip.

**Speaker notes**

The cluster bootstrap resamples query identities because the target is an equal-query mean. Duplicating a selected query duplicates its attached statistic. The exact pseudo-random generator, seed, resample count, and endpoint convention belong in the manifest. A row bootstrap is a different procedure.

**Teaching check**

Ask why the repeated response rows are not sampled independently in the primary route.

**Alt text**

“Twenty query cards are sampled with replacement by a seeded process to generate 5,000 equal-query mean contrasts.”

## Slide 15 — Seeded bootstrap result

**On-screen text**

Equal-query mention contrast:

`mean(SYN-B − SYN-A) = 0.084027778`

5,000-resample query-cluster percentile interval:

`[0.010416667, 0.159027778]`

Fixture-resampling result, not causal effect.

**Visual specification**

A compact histogram-like strip of bootstrap means uses an overlaid numeric endpoint bracket. A vertical line marks the point estimate. A hatched caption box lists no live engine, no randomized treatment, and no population frame.

**Speaker notes**

Read every decimal with its algorithm. Positive endpoints do not establish surface causality because surfaces were not assigned treatments and the fixture is authored. The interval represents resampling of twenty synthetic query clusters under an exact procedure; it does not cure query-frame bias or time drift.

**Teaching check**

Learners write one supported interpretation and two prohibited interpretations of the interval.

**Alt text**

“A seeded query-cluster bootstrap interval from 0.010416667 to 0.159027778 surrounds a point estimate of 0.084027778.”

## Slide 16 — Missing states stay visible

**On-screen text**

L06 statuses:

- complete: 354
- missing: 6
- refused: 0

Real protocols may add timeout · interface unavailable · parser failure · policy block.

Missing ≠ negative outcome.

**Visual specification**

Three status bins show counts. Six missing event cards retain their query, surface, block, and repeat keys while outcome cells remain blank. Additional real-world status labels appear as dashed schema extensions.

**Speaker notes**

A zero is an observed negative event. A blank after collection failure is unobserved. An absent citation interface may make citation undefined. A refusal is its own response state. Preserve designed cells and classify mechanisms before deciding on an outcome-specific denominator.

**Teaching check**

Ask how a parser failure differs from a complete response with mention zero.

**Alt text**

“Six missing rows retain full design keys and blank outcomes, distinct from 354 complete rows and any observed zero.”

## Slide 17 — Missing-policy sensitivity

**On-screen text**

Observed mentions across all blocks:

`SYN-A: 77/177`  
`SYN-B: 91/177`

Equal-query complete-case contrast: `0.084027778`  
Designed-event compound contrast: `(91−77)/180 = 0.077777778`

Different estimands.

**Visual specification**

Two parallel denominator paths begin from the same panel. One conditions on complete rows; the other uses all designed rows and renames the event response-and-mention. Their endpoints are printed side by side without a winner badge.

**Speaker notes**

The designed-event rule treats noncompletion as no observed response-and-mention. That can be useful operationally but is not mention incidence among returned responses. The equal-query complete-case estimate also changes weighting. Sensitivity should reveal these choices rather than hide them. Neither is a universal correction.

**Teaching check**

Learners supply a valid name for the designed-event compound metric.

**Alt text**

“Complete-case and designed-event paths yield 0.084027778 and 0.077777778 under explicitly different denominators for the frozen fixture.”

## Slide 18 — Query-weight sensitivity

**On-screen text**

SYN-A citation:

| Weighting | Estimate |
|---|---:|
| Equal strata | 0.356944 |
| Target weights 0.20/0.50/0.20/0.10 | 0.330556 |

Weights are authored sensitivity values, not observed demand.

**Visual specification**

Four stratum bars show their citation means. Two weight rows beneath produce two labeled totals. A dashed arrow from the weight source reads “declare before outcomes.”

**Speaker notes**

The target weights emphasize learn queries, where SYN-A citation is lower in the fixture. The estimate consequently falls. This is not a correction to truth; it answers a different target-distribution question. A report should show stratum values so weighting cannot hide heterogeneity.

**Teaching check**

Ask why weights chosen to maximize the contrast are analytically invalid.

**Alt text**

“Four stratum means combine under equal and unequal weights to produce SYN-A citation estimates of 0.356944 and 0.330556.”

## Slide 19 — Sensitivity is finite and preregistered

**On-screen text**

Predefine:

- query weights
- missing-state rule
- source/entity resolution
- incident-block inclusion
- time window
- cluster versus row resampling

Report sign, material size, and decision changes.

**Visual specification**

A finite six-row sensitivity grid has primary and alternate columns. Each row displays an outcome status: stable, materially changed, sign changed, or non-comparable. No open-ended tuning slider appears.

**Speaker notes**

Sensitivity evaluates dependence on justified assumptions. It is not a search across specifications for a favorable number. Freeze the finite set before comparative outputs. If a choice changes the estimand, name the alternate metric rather than presenting it as mere robustness.

**Teaching check**

Learners distinguish sensitivity to alias resolution from post-result cherry-picking of query weights.

**Alt text**

“A preregistered six-row sensitivity grid reports whether each justified alternative changes sign, size, decision, or comparability.”

## Slide 20 — Non-comparable metrics

**On-screen text**

Before comparing, align:

`event · unit · eligibility · query frame · surface · time · missingness · aggregation · uncertainty`

Matching labels are insufficient.

Output may be: **not directly comparable**.

**Visual specification**

Two metric cards enter a nine-field alignment checker. Several fields fail with written mismatch reasons. The checker routes to a “not directly comparable” box instead of forcing a converted score.

**Speaker notes**

Comparability is earned field by field. A dashboard citation count can differ from a binary response-level citation incidence in sampling, identity resolution, and time. A paper’s absorption measure is not a dashboard citation count. The scientifically correct result can be non-comparability.

**Teaching check**

Ask which fields must match before pooling estimates from two studies that both say “visibility rate.”

**Alt text**

“Two metric cards fail a nine-field comparability check and are explicitly labeled not directly comparable.”

## Slide 21 — PLAT-04 as an interface case

**On-screen text**

PLAT-04 local catalog role:

- official Microsoft/Bing public-preview guidance
- citations, cited pages, sampled grounding queries, trends
- specified Microsoft surfaces

Not established by counts: rank · authority · placement · cross-engine share · business outcome.

**Visual specification**

An original documentation card lists named interface fields. A thick boundary separates documented fields from five hatched non-inferences. No brand logo, screenshot, or copied interface layout appears.

**Speaker notes**

Use PLAT-04 to practice claim-relative authority. Official documentation can define fields and scope for its named preview. It does not make those counts interchangeable with L06 events or reveal rank. The interface may also evolve, so version date and surface coverage belong in the record.

**Teaching check**

Learners write one claim PLAT-04 can support and one cross-engine claim it cannot.

**Alt text**

“A product-specific documentation card lists reported fields while rank, authority, placement, cross-engine share, and outcomes remain outside its boundary.”

## Slide 22 — Preregistered panel analysis

**On-screen text**

Freeze before comparative outputs:

question · panel · metric cards · weights · blocks · retry · missing policy · aggregation · cluster · interval · seed · sensitivities · decision · ceiling

Log every deviation.

**Visual specification**

A fourteen-line registration sheet receives a timestamped freeze seal. A separate deviation ledger branches from it, retaining original and amended plans rather than overwriting either.

**Speaker notes**

Preregistration aligns analysis with a target rather than a preferred result. A collection incident, parser revision, or interface change can require amendment. Preserve the original and explain consequences. The plan is not proof that assumptions are correct; it makes them inspectable.

**Teaching check**

Ask what to do if an outage occurs in B-02 after the plan is frozen.

**Alt text**

“A fourteen-field analysis plan is frozen before results, with later deviations preserved in a separate dated ledger.”

## Slide 23 — Reading and evidence ceilings

**On-screen text**

- **PAPER-29:** repeated and longitudinal measurement route
- **PAPER-33:** citation-selection and absorption measurement route
- **PAPER-38:** at-scale measurement extension
- **PLAT-04:** named-interface case
- **R07:** practitioner synthesis of two time scales

Paper status: preprint / venue not confirmed here.

**Visual specification**

Five source cards sit below individually labeled ceiling bars. Above them are blocked claims: universal repeat count, cross-engine equivalence, causal platform effect, and current hidden mechanism.

**Speaker notes**

Read each source for its actual setting and estimand. PAPER-29 does not supply a universal repetition count. PAPER-33’s constructs do not collapse into one dashboard score. PAPER-38 retains its sampling boundary. R07 organizes questions but does not replace primary validation. PLAT-04 remains product-specific.

**Teaching check**

Learners select one source and state its supported role and strongest prohibited generalization.

**Alt text**

“Five reading cards remain under separate evidence ceilings that block universal repeat, equivalence, causal, and hidden-mechanism claims.”

## Slide 24 — The bounded repeatability conclusion

**On-screen text**

Name:

1. event and denominator
2. query and analysis units
3. surfaces and blocks
4. repetitions and dependence
5. estimate and interval method
6. missing/weight sensitivity
7. within-block variation versus drift
8. validity ceiling

Exit: “Repeatable for what target, under which design?”

**Visual specification**

Eight sentence blocks form a conclusion template. Two parallel arrows keep within-block instability and between-block drift separate until a final bounded summary. A hatched outer frame lists live and causal claims as not evaluated.

**Speaker notes**

Close with the fixture result: 360 designed events, 354 complete, six missing; SYN-A mention declines descriptively across blocks while many cells disagree across repeats. The query-cluster interval follows an exact seeded algorithm. None of this establishes live engine behavior or cause. Precision is a property of the declared target and design.

**Teaching check**

Learners submit a two-sentence result containing a full metric card, one uncertainty statement, and the strongest unknown.

**Alt text**

“An eight-part conclusion template keeps repeated-run variation and time-block drift separate inside an explicit validity boundary.”
