Explain

W07 · complete online lecture

Metrics, Uncertainty, and Repeated Observation

Essential question

When is a visibility change repeatable?

A structurally complete authored draft

Structurally complete authored package draft · human review and timed pilot pending

3,143online lecture words
4,112transcript words
24specified slides
90+90scheduled contact minutes

What the package must enable

  1. Specify event, numerator, denominator, unit, query frame, weighting, time window, repetitions, missingness, aggregation, and uncertainty in a metric card.
  2. Keep mention, citation, entailment, absorption, referral, and user action as distinct event namespaces.
  3. Separate within-block repeated-run variation from between-block descriptive drift.
  4. Resample at the declared query-cluster unit and interpret intervals under their dependence ceiling.
  5. Evaluate predefined weighting and missing-state sensitivities without post-result selection.

Planned 90-minute evidence sequence

W07 seminar plan
MinutesSegmentLearner evidence
0–19 minDenominator audit and event ladderCorrected metric card plus event-specific units
19–43 minPanel anatomy and two time scalesComplete cell key and separate variance/drift summaries
43–66 minDependence, clustering, Wilson intervals, and bootstrapCorrect resampling unit and bounded interval statement
66–86 minWeighting, missingness, and interface comparabilitySensitivity table and product-specific boundary memo
86–90 minExit testRepeatability conclusion and strongest unknown

75-minute core artifact route inside a 90-minute studio

The remaining 15 minutes are a declared delivery margin for setup, accessible pacing, questions, recovery, and submission packaging; they are not unplanned teaching content.

W07 core studio plan
MinutesActivityStop or redirect condition
0–18 minFreeze the analysis plan and validate the 360-cell L06 designReject ambiguous outcomes, denominators, or silently absent designed cells.
18–41 minReproduce rates/intervals and separate the two time scalesDo not treat row-level Wilson intervals as dependence-adjusted inference.
41–65 minRun the seeded query-cluster bootstrap and sensitivity analysisReject row bootstrap as the primary population interval and retain predeclared rules.
65–75 minAudit comparability and package a bounded resultBlock live, causal, or population-generalization claims from the fixture.

Read, teach, inspect, or download

Complete lecture text

This HTML is generated from the controlled Markdown source. Source SHA-256: 3ce7269f5eea1ff6fcb24151136a0b1c808df6a51349e850f6479347b6806e11.

On this page 11 sections

1. A metric name is not a measurement specification

“Visibility rose by twelve percent” sounds quantitative but may say almost nothing. What event rose: entity mention, source citation, claim entailment, source-distinctive absorption, referral, or a user action? Twelve percent of which eligible cells? Was the change twelve percentage points or twelve percent relative to baseline? Were missing responses removed, scored zero, or preserved as a separate outcome? Were twenty query intents sampled once or one query repeated twenty times? Without these fields, the number cannot be reconstructed or compared.

W07 uses a metric card with at least fourteen entries:

  1. metric ID and version;
  2. target event and decision rule;
  3. numerator;
  4. eligible denominator;
  5. observation and analysis units;
  6. query population and strata;
  7. query weights;
  8. surfaces, locales, and configuration;
  9. time window and time blocks;
  10. repetitions and retry rule;
  11. missing-state categories and outcome-specific policy;
  12. aggregation order;
  13. uncertainty method and resampling/model unit;
  14. interpretation ceiling and decision rule.

The card is written before the headline estimate. It prevents an analyst from switching denominators or weights after seeing which version produces a preferred narrative. A change to any field creates a new metric version, even when the display label stays the same.

PAPER-29 motivates repeated and longitudinal observation within its own study design. PAPER-33 offers a route for separating citation selection, support, and absorption. PAPER-38 extends measurement-at-scale questions. Their local status is preprint / venue not confirmed. A shared word such as “visibility” does not make their samples, surfaces, outcomes, or time windows comparable. PLAT-04 is a product-specific public-preview interface case, not a universal definition service. R07 is a practitioner synthesis that motivates two time scales but remains below primary empirical evidence in the source hierarchy.

Checkpoint A. Rewrite “citation rate is 40 percent.” A complete answer might say: among complete query–surface–time–repetition responses in the frozen panel, 40 percent contained at least one resolved citation event under rule CITE-v1; noncomplete responses remained explicit and were excluded from the outcome denominator; estimates were reported by surface and time block.

2. Event ladders prevent metric substitution

Six events can occur in sequence without being equivalent:

  • mention: a resolved entity appears in a response;
  • citation: a resolved source object is visibly attributed;
  • entailment: an exact cited passage supports an adjacent atomic claim;
  • absorption: source-distinctive, verified information appears faithfully in the answer under a declared rule;
  • referral: a recorded user or system event reaches a source;
  • action: a qualified downstream task event occurs.

Each event changes the unit. Mention may be entity–response. Citation may be source–response. Entailment is claim–passage–response. Absorption requires a preregistered source-distinctive claim set. Referral and action are user–task events with exposure and privacy requirements. Dividing each numerator by “number of prompts” does not erase those differences.

Conditioning can create additional incompatibility. Citation correctness among successfully fetched cited pages estimates a fetch-conditioned object. Citation incidence among all complete responses estimates another. Referral among citation-bearing answers estimates a post-citation subset. The denominator flow must retain planned cells, returned responses, citation-bearing responses, resolved citations, fetched pages, segmented claims, and adjudicated pairs as relevant.

Composites are especially risky. Averaging mention, citation, and referral produces a number with no natural unit unless a decision problem and weights justify it. An improvement in mention can hide a decline in entailment. W07 therefore reports the event ladder separately before any composite is proposed.

PLAT-04 illustrates the comparability audit. According to the local catalog, its preview reports citations, cited pages, sampled grounding queries, and trends on specified Microsoft surfaces. Those interface objects have documented product meanings. They do not automatically equal a course citation-event row, a source rank, authority, placement, cross-engine share, or user action. A dashboard label is evidence about the named interface documentation, not a transferable metric contract.

Checkpoint B. Can a citation count be compared with an absorption rate? Not until both events, units, denominators, source resolution, claim set, time window, and aggregation are aligned. Even then, they remain different outcomes rather than interchangeable measures.

3. Build the panel before collecting outcomes

An evaluation cell should identify query, surface, observable configuration, locale, time block, repetition, and collection mode. L06 uses the key:

query_id × surface × time_block × repetition.

Its frozen design contains 20 queries, two synthetic teaching surfaces, three blocks, and three repetitions. The Cartesian product has 360 designed cells. Every combination remains a row, including six noncomplete events. A silently absent row cannot be distinguished from a design omission, collection failure, or data-processing loss.

The query bank needs a target interpretation. L06 has four intent strata—compare, learn, select, and verify—with five queries each. Equal stratum counts make diagnostics convenient. They do not prove that real demand assigns each stratum 25 percent weight. A target-weighted estimate requires weights declared from a sampling design, decision population, or independently estimated usage. Each source defines a different estimand.

Panel completeness is checked before outcomes. Validate unique event IDs, unique combination keys, consistent query-to-stratum mapping, allowed statuses, timestamps with offsets, surface and block identities, repetition ranges, and the synthetic marker. Hash the panel. Freeze the analysis plan. Then compute estimates.

Retain raw event rows after aggregation. A surface-by-block table cannot reveal whether instability is concentrated in one query or evenly distributed. It cannot show whether one query has two complete repetitions while another has three. The archive supports alternate estimators and incident review.

Retries need a rule. If a missing response triggers repeated attempts until success, the observed dataset conditions on eventual return and hides failure probability. A protocol might allow one technical retry while preserving the original cell and linking the retry as a new event. It must not silently replace the first record.

Checkpoint C. A designed cell never appears in the CSV. Should it be coded zero? No. First classify whether the design, collection, ingestion, or validation process failed. Zero is an observed negative outcome, not a generic missing marker.

4. Numerators, denominators, and aggregation order

For a binary event, the simplest complete-case rate is successes / complete outcome observations. L06 applies this rule separately for each metric within each surface and time block. For SYN-A in block B-01, mention has numerator 28 and denominator 58, producing 0.482759. The designed denominator is 60, but two response rows are missing. Treating those missing responses as negative would produce 28/60, a different quantity.

Aggregation order determines weights. A row-level complete-case rate weights each observed repetition equally. An equal-query estimator first averages each query’s available repetitions, then averages the 20 query means. If some queries have more missing rows, these estimators differ. A stratum estimator can first compute query means within each intent and then apply target weights. None is automatically correct; the metric card selects the target.

Weights should sum to one and belong to named strata or queries. Equal diagnostic weights answer “what is the mean across this balanced teaching panel?” Usage weights answer “what is expected under this declared demand distribution?” Business priorities can define a decision-weighted score, but it should not be mislabeled as population prevalence.

Relative and absolute changes also differ. A rate change from 0.20 to 0.25 is +0.05 percentage points and +25 percent relative to baseline. Report both only when each is useful and label the denominator of the relative change. When baseline is zero or very small, relative change can be undefined or unstable.

An event-specific denominator can be conditional. Entailment among adjudicated citation–claim pairs is not the same as entailment incidence among all responses. A denominator funnel keeps the conditioning explicit. The chosen analysis must not convert one into the other by dropping intermediate failures.

Checkpoint D. Why can two analysts use the same 360-row file and obtain different valid estimates? They may target complete response rows, designed cells, equal query means, or weighted strata. The estimates answer different questions if the choices are declared; they become misleading when the choice is hidden.

5. Two time scales: repetition and drift

Repeated runs within one query–surface–block cell describe instability under nominally similar conditions. For binary outcomes, a cell with values [1, 1, 0] has mean two thirds and nonzero within-cell sample variance. More repeats refine the description of that cell’s run distribution. They do not add a new query intent or later platform state.

Between-block change describes a different time scale. L06 reports each surface’s last-block rate minus its first-block rate. For SYN-A mention, the complete-case rate moves from 0.482759 in B-01 to 0.400000 in B-03, a difference of −0.082759. This is descriptive drift in the synthetic fixture. It does not identify why the authored panel changed.

Do not compare within-cell variance and a between-block rate difference as if they had the same unit. A useful report places them beside one another: proportion of query cells showing run disagreement within each block; distribution of query-level cell means; first-to-last block rate; and individual query trajectories. The two summaries answer whether the result varies among close repetitions and whether the panel state differs across blocks.

Pooling blocks can hide drift. If a surface improves then declines, an overall mean may resemble baseline while the time path is operationally important. Conversely, a small block delta does not prove stationarity. Three blocks provide a teaching comparison, not a general time-series model.

PAPER-29 is assigned because repeated measurement is a first-class empirical problem. The course does not copy its reported repeat count as a universal prescription. Required repetitions depend on the event rate, query heterogeneity, surface behavior, time window, missingness, desired precision, and cost.

Checkpoint E. One query is repeated ten times in one hour. Does that establish stability across a month? No. It estimates one within-window component and supplies no later time block.

6. Repeated rows are dependent observations

Events sharing a query, surface, and block are not independent draws from the target query population. They share wording, intent, configuration, and time state. Treating 360 rows as 360 independent query intents exaggerates information about query-population coverage.

Choose the independent sampling or assignment unit. If the target is the mean across 20 sampled query intents, a nonparametric bootstrap should resample query IDs and retain all attached surface, block, repetition, and status rows. A row bootstrap represents an exchangeability assumption across response events. It answers a different and usually narrower uncertainty question.

The intraclass correlation coefficient can describe within-cluster similarity under a model. A design-effect expression such as 1 + (m - 1) rho illustrates how positive correlation reduces independent information from additional rows. It is not a universal correction. Unequal clusters, crossed surface/time factors, missingness, and interactions require more careful methods.

There are at least three defensible analysis levels:

  1. cluster summaries, such as one mean per query, followed by transparent comparisons;
  2. cluster bootstrap or randomization inference aligned with sampling or assignment;
  3. a preregistered hierarchical model that represents query, surface, block, and interactions.

W07 uses the first two for inspectability. It does not fit a causal model and does not treat surfaces as randomized treatments.

Increasing repetitions can reduce uncertainty about run variation within existing query clusters. It cannot repair a panel with too few query intents, an unrepresentative query universe, or missing time states. Thirty repeats of one favorable prompt remain one prompt cluster.

Checkpoint F. Twenty queries each have eighteen rows. What is the primary cluster count for an equal-query population estimate? Twenty, not 360. The other rows provide repeated measurements within those clusters.

7. Intervals: method, target, and ceiling

L06 reports 95 percent Wilson score intervals for each binary surface–block metric using complete-response denominators. The Wilson calculation behaves better than a simple normal approximation for some small samples and rates near zero or one. It still treats the supplied successes and trials according to a binomial description. In this repeated panel, it is descriptive and not dependence-adjusted.

Interval overlap is not a hypothesis test. Two overlapping intervals can accompany a detectable paired difference; two nonoverlapping marginal intervals do not replace a prespecified comparison model. A narrow interval around a biased estimator is precisely wrong. Report the sampling frame, dependence, missingness, and measurement error beside numeric endpoints.

The W07 cluster-bootstrap demonstration targets the mean across the 20 query-specific SYN-B minus SYN-A mention differences. For each query, calculate the mean among complete observations for each surface across all blocks and repetitions, then subtract. The point estimate is 0.084027778. With seed 707, 5,000 resamples of 20 query IDs, and a declared empirical-percentile index rule, the synthetic interval is [0.010416667, 0.159027778].

This interval describes how that statistic varies when the 20 fixture query clusters are resampled under the stated algorithm. It does not correct query-frame bias, establish surface causality, represent live engines, or model time drift. Resampling all attached rows preserves within-query structure but also carries the frozen fixture assumptions.

Bootstrap records include cluster definition, number of resamples, random generator, seed, statistic, stratification, missing policy, and percentile method. Different choices can change endpoints. Do not choose a seed or interval type after observing the desired conclusion.

Checkpoint G. What does the L06 Wilson interval omit? At minimum, it does not adjust for repeated-query dependence, query-frame uncertainty, label error, or drift mechanisms.

8. Missing is a state, not a negative outcome

L06 permits complete, missing, and refused. Complete rows contain binary outcome values. Noncomplete rows retain blank outcomes. The fixture happens to contain six missing and no refused events, but the schema preserves both categories.

A real protocol may need more states: timeout, unavailable search, absent citation interface, parser failure, blocked collection, policy refusal, authentication change, redaction, or corrupted artifact. These mechanisms imply different next actions. An absent citation interface makes citation unobservable, not necessarily zero. A refusal is an answer outcome of one kind while downstream mention or citation fields may be undefined. A parser failure is a measurement failure.

Missingness analysis starts with designed, returned, missing, and refused counts by surface, block, query stratum, and relevant configuration. Then inspect whether noncompletion depends on query, surface, time, or observed outcomes. Complete-case estimates target returned responses under assumptions. Designed-denominator estimates that score noncomplete events zero target a compound “response-and-event” quantity. Both can be reported as a sensitivity when clearly named.

In L06, both surfaces have 180 designed events and 177 complete events. SYN-A has 77 observed mentions; SYN-B has 91. The designed-denominator difference is (91 - 77) / 180 = 0.077777778. The equal-query complete-case difference is 0.084027778. These numbers differ because aggregation and noncompletion handling differ. Neither should replace the other silently.

Worst-case bounds can assign missing outcomes in directions that challenge a conclusion. With few symmetric missing cells, bounds may be narrow; with differential noncompletion, they can dominate the result. A model-based imputation adds assumptions and should not be the only analysis.

Checkpoint H. A response artifact is missing after a collection error. Is mention zero? No. Mention is unobserved. A separate compound metric may score failure as no observed mention, but it requires a new name and denominator.

9. Query weighting and sensitivity analysis

The L06 panel is balanced across compare, learn, select, and verify intents. Its equal-stratum diagnostic mean answers a balanced-panel question. Suppose a declared target population assigns weights 0.20, 0.50, 0.20, and 0.10 respectively. Applying these weights before inspecting a headline creates another estimand.

For citation on SYN-A, equal-query stratum means are approximately 0.388889 for compare, 0.288889 for learn, 0.333333 for select, and 0.416667 for verify. Their equal-stratum mean is 0.356944. Under the declared target weights, the estimate is 0.330556. For SYN-B, the corresponding equal-stratum estimate is 0.454861 and target-weighted estimate is 0.460556. Query mix changes both levels and the contrast.

A sensitivity table should be finite and preregistered. Useful rows include:

  • equal query weights;
  • target demand weights from a declared source;
  • complete-case versus designed-event compound outcome;
  • primary source-resolution rule versus a stricter alias rule;
  • inclusion and exclusion of incident-affected blocks;
  • cluster versus row resampling, clearly labeled as different targets;
  • alternate but justified time windows.

Do not search across dozens of weights and publish the most dramatic. The table should show which decisions change the sign, material size, or decision category. Sensitivity is evidence about robustness to assumptions, not permission to select an estimate.

If two metrics are non-comparable, sensitivity cannot fix them. A PLAT-04 count and a PAPER-33 absorption measure may both concern sources but operate on different events and eligible populations. The correct outcome is “not directly comparable” with a field-by-field explanation.

10. Preregister the panel analysis

A W07 pre-analysis record is written before viewing comparative outputs. It names:

  • research question and descriptive estimand;
  • query universe, sample, strata, and weights;
  • surfaces/configurations and comparability requirements;
  • block dates or authored identities and repetitions;
  • primary and secondary metric cards;
  • missing categories, retry, exclusion, and incident rules;
  • primary aggregation and uncertainty;
  • cluster bootstrap seed, replicates, and percentile rule;
  • predefined sensitivity analyses;
  • decision thresholds and reporting of adverse results;
  • claim ceiling and generalization boundary.

Deviations remain visible. If an outage requires a new block, preserve the original plan and append a dated deviation. If a parser changes, reprocess all comparable cells or version the metric. If a platform interface changes, do not merge pre-change and post-change rows without a declared estimand.

L06 is deterministic because the panel and analysis are frozen, not because the underlying phenomenon represented by the teaching story is deterministic. Re-running the script should reproduce exact outputs and hashes. Reproducibility of code and fixture is different from repeatability of a live phenomenon and different again from external validity.

Ethics and operational constraints remain part of measurement. The core route uses no live calls. Any optional collection requires permission, terms/rate review, cost cap, privacy review, account-risk controls, and explicit provenance. More repeats do not justify unauthorized load.

11. Bounded conclusions, discussion, and exit test

A complete W07 conclusion contains:

  1. panel and event definition;
  2. query and analysis units;
  3. surface and block identities;
  4. numerator, denominator, weighting, and missing rule;
  5. repetition and dependence structure;
  6. estimate and interval method;
  7. within-block instability and between-block drift separately;
  8. predefined sensitivity results;
  9. non-comparable measures and strongest validity limitation;
  10. claims not made.

Example:

In the 360-event L06 synthetic panel, SYN-A complete-response mention incidence declined from 28/58 in B-01 to 24/60 in B-03, a descriptive last-minus-first difference of −0.082759. The panel also showed repeated-run disagreement within query–surface–block cells. These summaries describe separate time scales in the authored fixture; they do not identify a drift mechanism or a live platform effect.

Cluster-bootstrap example:

The equal-query mean of complete-case SYN-B minus SYN-A mention differences was 0.084027778 across 20 fixture query clusters. A seeded 5,000-resample query-cluster percentile procedure produced [0.010416667, 0.159027778] under its exact algorithm. The interval represents resampling of the frozen query set, not causal identification, demand-population coverage, or named-engine behavior.

Discussion prompts

  1. When is the designed denominator preferable to the complete-response denominator, and how should the metric be renamed?
  2. Why can more repetitions fail to improve query-population coverage?
  3. What would make two citation-rate labels non-comparable?
  4. How could missingness depend on an outcome?
  5. Which uncertainty procedure matches a mean over sampled query intents?
  6. What does PLAT-04 document, and what rank or placement conclusion remains unsupported?
  7. How do fixture reproducibility, empirical repeatability, and external validity differ?

Exit test

Write two sentences. The first must specify a metric card and a repeatability result with an interval or sensitivity. The second must distinguish within-block instability from between-block drift and state the strongest unknown. If any denominator, weighting, missing-state rule, cluster unit, or time block is recoverable only from oral explanation, revise the submission.

Frozen L06 Repeated-Observation Panel · synthetic from input to boundary

A 360-cell synthetic panel preserves six noncomplete cells, distinguishes repeated-run variation from block drift, and reproduces query-cluster, weighting, aggregation, and missing-state sensitivities under declared ceilings.

Controlled source route

Core PAPER-29/PAPER-33 · Extend PAPER-38 · PLAT-04 interface case · practitioner synthesis R07 · NIST statistical-method guidance · Core Notes measurement chapter · offline L06 panel.