# W07 No-Video Transcript — Metrics, Uncertainty, and Repeated Observation

## Status and use

No W07 recording exists. This is a readable and recordable no-video equivalent, not a transcript of a completed lecture. Chapter durations are **instructional planning budgets**. They have not been verified through rehearsal, recording, caption synchronization, or learner use. A future recording must be checked for real duration, mathematical voicing, speaker labels, captions, audio description, and navigable chapter boundaries.

All empirical-looking values come from the frozen L06 synthetic fixture. No live platform is queried or represented. The transcript describes every core visual in words and preserves the distinctions among event, denominator, cluster, time block, missing state, interval, sensitivity, and claim ceiling.

## Planned chapter budget

| Chapter | Planned minutes | Core purpose |
|---:|---:|---|
| 1 | 4 | Replace vague visibility with a metric card |
| 2 | 4 | Separate stage events and denominator flows |
| 3 | 4 | Build and validate the repeated panel |
| 4 | 4 | Separate within-block variation from drift |
| 5 | 4 | Preserve dependence and choose the cluster |
| 6 | 4 | Interpret Wilson and cluster-bootstrap intervals |
| 7 | 4 | Classify missing states and denominator sensitivities |
| 8 | 4 | Apply query weights and comparability audits |
| 9 | 5 | Preregister analysis and reproduce L06 |
| 10 | 5 | Apply source ceilings and write the conclusion |

**Total planned route: 42 minutes.**

## Chapter 1

Welcome to W07, metrics, uncertainty, and repeated observation. The essential question is: when is a visibility change repeatable?

Begin with the sentence “visibility increased by twelve percent.” It contains a number but not yet a measurement. We need to know the event. Was an entity mentioned? Was a source cited? Did a cited passage entail an answer claim? Did source-distinctive information appear? Did a user reach the source? Did a qualified action occur?

We also need a denominator. Twelve percent of designed query cells, returned responses, citation-bearing responses, resolved sources, fetched pages, adjudicated claim pairs, or exposed users? Was the change twelve percentage points or twelve percent relative to baseline? Were missing responses omitted, treated as zero, or preserved as separate states?

W07 solves this ambiguity with a metric card. The card records a metric ID and version, target event, numerator, eligible denominator, observation and analysis units, query frame and strata, query weights, surfaces and locale, time blocks, repetition and retry rule, missing categories, aggregation order, uncertainty method, and interpretation ceiling.

The card is frozen before comparative outputs. If the denominator or weight changes, create a new metric version. Do not retain the same display label and act as if the number remained comparable.

Here is a complete example. Among complete query–surface–time–repetition responses in the frozen L06 panel, mention incidence is the count of rows with a resolved entity-event label equal to one divided by complete rows eligible for that event. Noncomplete rows remain in the panel with blank outcomes. Estimates are reported by synthetic surface and authored block. The Wilson interval is descriptive and not dependence-adjusted.

Notice how much more limited this is than “visibility.” The precision is an advantage. A reader can reproduce the numerator and denominator and knows what was not measured.

PAPER-29 motivates repeated and longitudinal measurement in its research setting. PAPER-33 supplies a route for separating citation selection from support and absorption. PAPER-38 extends measurement-at-scale questions. The local catalog identifies all three as preprints whose venue status was not confirmed. Their shared topic does not make their estimates interchangeable.

The first checkpoint is to rewrite the vague sentence. Name one event, one eligible denominator, one panel, one time window, and one uncertainty method. If you cannot, you do not yet have a metric.

## Chapter 2

We now separate six events. Mention is the appearance of a resolved entity in a response. Citation is visible attribution to a resolved source object. Entailment asks whether an exact passage supports an adjacent atomic claim. Absorption asks whether verified, source-distinctive information appears faithfully under a declared rule. Referral is a user or system event reaching a source. Action is a qualified user–task outcome.

These events have different units. Mention may use entity–response records. Citation may use source–response records. Entailment uses claim–passage–response records. Absorption requires a preregistered claim set. Referral and action require exposure, users or tasks, event logging, and privacy controls.

Do not average them into one “visibility score” merely because each can be represented between zero and one. A number without a common unit and decision function is not made meaningful by normalization. Report the event ladder separately first.

Denominator flows make conditioning visible. Imagine designed cells, returned responses, citation-bearing responses, resolved citation objects, successfully fetched passages, and adjudicated claim–source pairs. A correctness rate among adjudicated pairs says nothing directly about citation incidence among all designed cells. Each narrowing step needs counts and reason codes.

A missing response is not a mention zero. A failed fetch is not an unsupported passage judgment. An unavailable citation interface is not necessarily no citation in an underlying system. We will return to missing mechanisms later.

PLAT-04 is useful for a field-definition audit. The local catalog describes official Microsoft/Bing public-preview guidance for an interface reporting citations, cited pages, sampled grounding queries, and trends across specified Microsoft surfaces. Those documented fields belong to that interface and version. The counts do not indicate source rank, authority, or placement. They are not automatically comparable with one synthetic L06 citation row or a PAPER-33 absorption construct.

The right comparison can be “not directly comparable.” Before comparing two rates, align event, unit, eligibility, query frame, surface scope, time, missingness, aggregation, and uncertainty. Matching labels are insufficient.

Here is the chapter check. Can citation count replace absorption? No. Absorption requires source-distinctive claims and answer-use evaluation. A visible citation event alone does not establish faithful use.

## Chapter 3

Next, build the panel before looking at outcomes. L06 uses one row for every query ID, synthetic surface, authored time block, and repetition. There are twenty queries, two surfaces, three blocks, and three repeats. Multiply them: twenty times two times three times three equals three hundred sixty designed events.

The combination key is query ID, surface, time block, and repetition. Every key must be unique. Every designed key must exist even when the response is missing or refused. Otherwise, absence in the file could mean design omission, collection failure, or processing loss.

The fixture has four query intent strata: compare, learn, select, and verify. Five queries belong to each. Equal numbers are useful for diagnostics. They do not prove equal prevalence in a target demand population.

Validate identities before outcome analysis. Event IDs are unique. Query-to-stratum assignments remain stable. Timestamps include offsets. Repetitions are positive integers. Status values are allowed. The synthetic marker is true. Complete rows have binary outcomes. Noncomplete rows have blank outcomes.

The panel hash is part of the contract. Re-running an analysis on a modified panel is a new version, even if the filename stays the same. Preserve the raw rows after aggregation because surface–block summaries cannot reveal individual query instability or missingness concentration.

Retries also belong in the design. Repeating until a response succeeds and then replacing the missing event conditions on eventual return. A defensible protocol can preserve the original failure and link a permitted retry as a new event. It should never erase the designed cell.

L06 contains three hundred fifty-four complete events and six missing events. There are no refused rows in this fixture, but the schema retains that state. Each surface has one hundred eighty designed and one hundred seventy-seven complete events.

The primary low-compute route is offline. It uses the frozen CSV and Python standard library. It requires no browser automation, paid account, API call, model download, or network request. This matters ethically and methodologically: the course can teach panel analysis without generating unauthorized load or pretending that synthetic patterns are live observations.

Checkpoint: if one designed event is not in the file, should it be coded zero? No. First locate the failure in design, collection, ingestion, or validation. Zero is an observed negative outcome.

## Chapter 4

We now separate the two time scales. Within one query–surface–block cell, repeated runs describe variation under nominally similar conditions. Between blocks, changes describe different authored time states.

Take query Q-001 on SYN-A in B-01. Its mention values are one, one, and zero. The cell mean is two thirds. The unbiased sample variance is one third. There is one query cluster and three repeated measurements.

Across the twenty SYN-A query cells in B-01, seventeen contain both zero and one among their complete repetitions. The mean within-query sample variance is approximately zero point two nine one seven. In B-02 and B-03, seventeen cells are also unstable, with mean sample variances approximately zero point two eight three three.

These are within-block summaries. They show that one run can misrepresent a cell. They do not describe month-scale drift or query-population uncertainty.

Now look across authored blocks. SYN-A complete-response mention is twenty-eight of fifty-eight in B-01, or zero point four eight two seven five nine. In B-03 it is twenty-four of sixty, or zero point four. Last minus first equals negative zero point zero eight two seven five nine.

For SYN-B, mention moves from thirty-one of fifty-eight, or zero point five three four four eight three, to twenty-nine of fifty-nine, or zero point four nine one five two five. The difference is negative zero point zero four two nine five eight.

These are descriptive drift values in the synthetic fixture. They do not identify a platform release, corpus update, model change, session state, or seasonal mechanism. No named platform appears in the data.

Do not compare sample variance and a block rate difference as though they had the same unit. Place separate summaries beside each other: cell disagreement, distribution of query cell means, block-level rates, first-to-last change, and query trajectories.

Pooling all blocks can hide the path. Repeating a single block many times cannot create later states. PAPER-29 helps motivate repeated measurement, but its repeat counts are properties of its setting, not a universal prescription. Official NIST uncertainty examples likewise warn that repeated observations may be time-dependent or otherwise non-independent; that is method guidance, not a GEO effect estimate.

Checkpoint: ten runs in one hour establish what? They can refine one within-window distribution. They do not establish stability over a month.

## Chapter 5

The panel contains three hundred sixty rows but not three hundred sixty independent query intents. Rows sharing a query also share wording, intent, and many surface or time conditions. Dependence affects how much independent information the rows provide.

If the target is a mean across twenty sampled query intents, use query as the primary cluster. One transparent estimator first averages each query’s complete rows for SYN-A and SYN-B, forms a within-query contrast, then averages the twenty contrasts.

A cluster bootstrap resamples query IDs with replacement and retains each selected query statistic or its attached rows. A row bootstrap resamples response events as if exchangeable. These procedures represent different uncertainty models. Row resampling can describe variation among observed events while omitting uncertainty about which queries entered the panel.

The intraclass correlation coefficient can describe similarity inside clusters under a specified model. A design-effect formula can illustrate why correlated repetitions add less independent information than new clusters. It is not a universal correction, especially with crossed surfaces and blocks, unequal complete counts, and interactions.

More repetitions can improve understanding of run-level variation in existing queries. They cannot repair an unrepresentative query universe. A sixth repeat does not replace five lost query intents. One hundred paraphrases of one fact need can still form one underlying intent cluster for sampling and evaluation.

Possible analysis routes include cluster summaries, cluster bootstrap, randomization inference when assignment supports it, or a preregistered hierarchical model. W07 uses inspectable cluster summaries and a deterministic cluster bootstrap. It does not fit a causal effect model and does not treat synthetic surfaces as randomized treatments.

The panel is crossed as well as nested. Each query appears on both surfaces and in every block, while repetitions sit within the query–surface–block cell. A hierarchical description can include query effects, surface effects, block effects, query-by-surface interactions, and surface-by-block interactions. Writing those terms down is useful even when we do not fit the model: it prevents the residual row from becoming a container for every source of variation. A model would still require distributional assumptions, diagnostics, convergence checks, and a target interpretation.

Cluster choice follows the question. If the target were change across independently sampled time periods, three authored blocks would be inadequate as a large-sample time cluster analysis. If collection sites were independently sampled, site might become another cluster level. There is no single universal “clustered standard error.” The data-generating and sampling structures determine which dependence matters.

The chapter check is numerical. Twenty queries each contribute eighteen designed rows. How many primary query clusters are available for an equal-query population summary? Twenty. The remaining structure helps describe repeated surface and time behavior inside those clusters.

## Chapter 6

Intervals require a method, target, and ceiling. L06 reports ninety-five percent Wilson score intervals for each binary event within a surface and block, using complete responses as trials.

For SYN-A mention in B-01, the estimate is twenty-eight divided by fifty-eight, or zero point four eight two seven five nine. The Wilson interval runs from zero point three five nine two eight two to zero point six zero eight three seven seven.

This is useful descriptive information for the declared binary denominator. It is not dependence-adjusted. It does not represent uncertainty from sampling the query frame, label error, missingness mechanism, or time drift.

Interval overlap is not a formal equality test. Nonoverlap is not a substitute for a prespecified paired or clustered comparison. A narrow interval around an estimator built on biased queries remains biased.

W07 adds a query-cluster bootstrap for the equal-query SYN-B minus SYN-A mention contrast. For each query, average complete mention rows across all blocks and repetitions for each surface. Subtract A from B. Then average the twenty query differences. The result is zero point zero eight four zero two seven seven seven eight.

Initialize Python’s standard pseudo-random generator with seed seven hundred seven. For each of five thousand resamples, sample twenty query indices with replacement and average the corresponding query differences. Sort the five thousand means. Use the declared empirical endpoint indices. The resulting interval is zero point zero one zero four one six six six seven to zero point one five nine zero two seven seven seven eight.

Positive endpoints do not establish causality. The synthetic surfaces were not assigned treatments, and the queries do not represent a documented demand population. The interval describes resampling of this frozen set under this algorithm. It also does not model the authored time trajectory.

Bootstrap documentation includes cluster, statistic, seed, generator, resample count, missing rule, stratification, and interval construction. NIST method guidance reinforces that bootstrap conclusions inherit sampling and dependence assumptions. It is not evidence about GEO behavior.

An interval report should show the point estimate, endpoints, confidence or coverage label, method, resampled unit, replicate count, and all material conditioning. It should avoid saying that there is a ninety-five percent probability that a fixed parameter lies inside this realized frequentist interval. The procedure’s interpretation concerns repeated applications under its assumptions. A percentile bootstrap also has known limitations in finite and skewed settings; W07 uses it because the algorithm is transparent, not because it is universally preferred.

When the number of independent clusters is small, an apparently smooth bootstrap distribution can overstate what the frame supports. Twenty synthetic query clusters are enough for this teaching computation, but not a certification of coverage. Report the individual query contrasts beside the interval so heterogeneity and influential clusters remain inspectable.

Checkpoint: what does the Wilson interval omit? Repeated-query dependence, query-frame coverage, label error, missing mechanism, and drift cause, among other things.

## Chapter 7

Missing is a state, not a negative event. In L06, complete responses have binary values for mention, citation, entailment, absorption, and referral. Missing or refused rows retain blank values.

A live protocol may need timeout, search unavailable, citation interface absent, parser failure, policy refusal, authentication change, blocked collection, redaction, and corrupted artifact. Each has a different mechanism and perhaps a different eligible denominator.

An unavailable citation interface makes citation unobservable under that collection route; it need not mean the source was absent from every internal representation. A parser failure is a measurement failure. A refusal is an observed response state, but downstream event labels may be undefined. Preserve this structure.

L06 has six missing rows and no refusals. Across all blocks, SYN-A has seventy-seven observed mentions among one hundred seventy-seven complete rows. SYN-B has ninety-one among one hundred seventy-seven.

The pooled complete-row difference is ninety-one divided by one hundred seventy-seven minus seventy-seven divided by one hundred seventy-seven, approximately zero point zero seven nine one. The equal-query complete-case contrast is zero point zero eight four zero three. They differ because aggregation and missing-cell locations affect weights.

Now define a compound operational metric: observed response-and-mention among all designed events. Treat a noncomplete event as no observed response-and-mention. The surface difference is ninety-one minus seventy-seven divided by one hundred eighty, or zero point zero seven seven seven eight.

This compound metric is not mention incidence among returned responses. It combines response availability and mention. Rename it and report it as sensitivity rather than silently replacing the primary denominator.

Worst-case bounds can assign missing outcomes in directions that challenge a result. Model-based imputation adds assumptions and should not be the sole route. Start with status counts by surface, block, and stratum and inspect whether noncompletion is differential.

Three common missingness labels help organize assumptions, but they are not automatically established from the data. Missing completely at random means noncompletion is unrelated to observed and unobserved outcome information. Missing at random conditions the mechanism on observed variables. Missing not at random allows dependence on the unobserved outcome. In a generative interface, refusal or timeout can plausibly depend on query content, surface state, or answer behavior, so a complete-case estimate may target a selected response population. Describe evidence for a mechanism instead of assigning a label by convenience.

Checkpoint: a collection error leaves no response file. Is mention zero? No. Mention is unobserved. A separate operational compound can score no observed response-and-mention, but it requires a different name.

## Chapter 8

Query weighting changes the target. L06 has equal numbers in four intent strata. An equal-stratum mean answers a balanced-panel question. It does not estimate real demand prevalence.

For SYN-A citation, equal-query stratum means are approximately zero point three eight eight nine for compare, zero point two eight eight nine for learn, zero point three three three three for select, and zero point four one six seven for verify. The equal-stratum mean is zero point three five six nine four four.

Now apply authored sensitivity weights: compare zero point two, learn zero point five, select zero point two, and verify zero point one. The target-weighted estimate is zero point three three zero five five six. It falls because learn receives more weight and has a lower fixture estimate.

For SYN-B, the equal-stratum citation estimate is zero point four five four eight six one. The target-weighted estimate is zero point four six zero five five six. The weighting changes both surface levels and their contrast.

These target weights are constructed for instruction. They are not observed usage. Weight sources might be a probability sampling design, independent usage data, or a decision-priority distribution. Each defines a different estimand.

Sensitivity analysis should be finite and preregistered. It can include equal versus target weights, complete-case versus designed-event compound outcomes, primary versus stricter source resolution, incident-block inclusion, alternate time windows, and cluster versus row resampling. Report whether sign, material size, or decision category changes.

Do not tune weights after viewing outcomes. Sensitivity is not permission to search for a headline. If an alternative changes the estimand, give it a new metric name.

Non-comparability remains a valid output. Before pooling two “citation rates,” align event, unit, eligible denominator, query frame, surface, time, identity resolution, missingness, aggregation, and uncertainty. If fields conflict, do not invent a conversion.

PLAT-04 interface counts, PAPER-33 absorption constructs, PAPER-29 repeated-observation estimates, PAPER-38 at-scale measures, and L06 binary events retain their own protocols. Similar vocabulary does not create one metric.

A comparability table should preserve the outcome when only some fields align. Two studies might use the same binary citation event but different query frames and dates; their definitions align while their population estimates remain non-poolable. Another pair might share queries and surfaces but use source-level versus response-level denominators. Rather than collapsing everything into comparable or incomparable, identify exactly which contrast is supported: definition comparison, descriptive side-by-side reporting, harmonized reanalysis, or no quantitative comparison.

## Chapter 9

Preregister the panel analysis before comparative outputs. The plan begins with a research question and descriptive estimand. It names the query universe, sample, strata, and weights. It names surfaces and comparability requirements, authored or real time blocks, repetitions, and retry rules.

Every primary and secondary metric receives a card. Missing categories, exclusions, incidents, and parser versions are frozen. Aggregation order and uncertainty method are explicit. For the cluster bootstrap, record query as cluster, seed seven hundred seven, five thousand resamples, the statistic, and endpoint rule.

The plan also lists finite sensitivity analyses, decision thresholds, adverse-result reporting, and the claim ceiling. Deviations are appended rather than overwriting the original. If an outage creates a new block, retain the plan and add a dated amendment. If a parser changes, reprocess comparable cells or version the metric.

Now reproduce L06. Verify the panel hash. Run the standard-library analysis with minimum queries equal to twenty. The pass line reports three hundred sixty events, twenty queries, two surfaces, and three blocks.

Inspect all outputs. `metric_estimates.csv` contains surface, block, metric, numerator, complete denominator, rate, and Wilson endpoints. `time_drift.csv` contains first and last block rates and differences. `query_mix_sensitivity.csv` retains four strata. `missingness.csv` retains designed, complete, missing, and refused counts. The validity boundary says repeated events for one query are not independent and prohibits collapsing mention into citation, citation into entailment, entailment into absorption, or absorption into referral.

Reproducibility means the same frozen inputs and script create the same outputs and hashes. Empirical repeatability means a phenomenon shows a defined stability pattern under repeated observation. External validity asks whether results transfer to another population or system. These are different achievements.

The core route makes no live calls. Optional collection would require permission, terms and rate review, cost cap, privacy controls, collection provenance, and account-risk safeguards. Additional repetitions never authorize otherwise prohibited collection.

## Chapter 10

We finish by applying evidence ceilings and writing a bounded conclusion.

PAPER-29 supports questions about repeated and longitudinal measurement in its published setting. It does not supply a universal repeat count. PAPER-33 supports distinctions among citation selection, support, and absorption in its setting. It does not make those measures identical to a product dashboard. PAPER-38 extends at-scale measurement questions while retaining its sampling and platform boundary.

PLAT-04 is official for the named Microsoft/Bing public-preview documentation and version. It can establish documented interface fields and scope. Its aggregate counts do not reveal rank, authority, placement, cross-engine share, or business outcomes. R07 is a practitioner synthesis used to organize within-run variance and between-time drift; it does not replace the primary evidence it synthesizes. NIST uncertainty material supports method caution about non-independent and time-dependent observations; it does not report a GEO empirical effect.

A complete W07 result names event, denominator, query and analysis units, weights, surfaces, time blocks, repetitions, dependence, missing policy, estimate, interval method, sensitivity, and validity ceiling.

Here is the frozen-panel result. In the three-hundred-sixty-event L06 synthetic panel, SYN-A complete-response mention incidence declined from twenty-eight of fifty-eight in B-01 to twenty-four of sixty in B-03, a descriptive difference of negative zero point zero eight two seven five nine. Many query–surface–block cells also disagreed across repeated runs. These summaries describe two distinct time scales in the authored fixture; they do not identify a mechanism or a live platform effect.

Here is the bootstrap result. Across twenty fixture query clusters, the equal-query complete-case mean of SYN-B minus SYN-A mention was zero point zero eight four zero two seven seven seven eight. An exact five-thousand-resample query-cluster procedure with seed seven hundred seven produced an empirical interval from zero point zero one zero four one six six six seven to zero point one five nine zero two seven seven seven eight. It represents fixture-query resampling under the stated algorithm, not causal identification or demand-population coverage.

The exit question is: repeatable for what target, under which design? Submit two sentences. The first must state a complete metric card and interval or sensitivity. The second must separate within-block instability from between-block drift and name the strongest unknown. If a reader needs oral clarification to recover the denominator, weights, missing policy, cluster, or time blocks, revise the claim.
