# W06 Slide Script — Queries as a Measurement Instrument

## Production conventions

This script specifies exactly 24 slides for an original academic deck. Use deep purple for scope and structure, gold for declared design decisions, charcoal for observed records, and a red hatch only for exclusions or unsupported transfer. Color must never be the sole encoding. Use labels, patterns, position, and line style. All diagrams are newly designed for W06; do not import figures, screenshots, or decorative fragments from papers, platforms, the attached TeX source, or other courses. Body text must remain readable at seminar distance, equations require a spoken reading, and every data slide needs a text table in the learner handout.

## Slide 01 — The prompt list is the instrument

**On-screen text**

> W06 · Queries as a Measurement Instrument  
> Essential question: What distribution do the prompts represent?  
> Design before collection.

**Visual specification**

Place a narrow stack of identical prompt cards on the left and a much larger field of differently shaped task tokens on the right. Between them, draw a labeled lens: “sampling design.” A dotted ray reaches only part of the field. The uncovered region stays visible. Add a bottom caption: “Row count is not population coverage.” Use no brand mark or platform UI.

**Speaker notes**

Open by asking whether 500 prompts are necessarily more representative than 50. The correct answer is that size alone does not resolve what the prompts were intended to represent, how they entered the list, or which units could never enter. A visibility estimate is conditional on the query instrument. If every prompt names the target, the result concerns prompted recognition. If every prompt asks for a recommendation, it does not describe factual lookup. W06 is about constructing the object that later metrics average over. The lesson deliberately precedes response collection: a favorable outcome cannot be allowed to choose the population after the fact.

**Teaching check**

Ask learners to complete: “Five hundred prompts establish ___, but do not establish ___.” Accept “a nominal row count” and “coverage or representativeness.”

**Alt text**

A lens labeled sampling design maps a small card stack to only part of a varied task population; a large uncovered region shows that a long prompt list can miss much of the declared universe.

## Slide 02 — Outcomes and evidence artifacts

**On-screen text**

- Define universe, frame, sample, and cell.
- Freeze strata, rules, provenance, clusters, and splits.
- Govern qrels and leakage.
- State external validity.

Artifacts: frame · query bank · codebook · exclusions · duplicate audit · label protocol · coverage map · hashes

**Visual specification**

Create an eight-spoke wheel around a central circle labeled “versioned instrument.” Each spoke ends in one artifact name. Use alternating solid and outlined nodes so the diagram remains legible in grayscale. Add a small gate below: “answers collected only after freeze.”

**Speaker notes**

Explain that the deliverable is not a clever list of prompts. It is a governed package another analyst can inspect without oral explanation. The frame says what may enter. The query bank preserves exact records and stable IDs. The codebook defines labels. Exclusions show the denominator loss. The duplicate report exposes dependence candidates. The label protocol controls qrels and other judgments. The coverage map states absences. Hashes make the frozen version identifiable. Learners will use the L02 fixture to rehearse this package, then build a larger 60-query instrument in the full assignment. The 20-row fixture is intentionally too small and should retain its warning.

**Teaching check**

Ask which artifact prevents excluded unfavorable prompts from disappearing. Expected answer: the versioned exclusion log, tied to stable IDs and the frozen frame.

**Alt text**

Eight labeled artifacts surround a central versioned instrument, with a gate stating that answer collection begins only after the design is frozen.

## Slide 03 — A list implies a population claim

**On-screen text**

“80 prompts” ≠ “the demand space”

Four questions:

1. Who asks?
2. What task?
3. Under which locale, time, and surface?
4. How could a unit enter?

**Visual specification**

Use a two-column proof ledger. The left column contains three progressively richer descriptions: “80 prompts,” “80 English comparison prompts,” and a bounded population statement. The right column shows unresolved fields as open boxes. Only the third row closes actor, task, entity, locale, time, and source; “coverage verified?” remains open.

**Speaker notes**

Move from metadata to design. “English” and “comparison” are useful labels, but still do not say whether prompts model shoppers, researchers, procurement teams, or synthetic test writers. A population statement should name actor or decision role, task family, entity scope, language and locale, surface or access mode, time window, and explicit exclusions. Even a complete statement is a claim to be audited against the frame. It does not certify that the operational frame reaches every declared unit. The framing sentence and the entry mechanism must be reviewed together.

**Teaching check**

Show “English questions about project-management tools” and ask for two missing dimensions. Accept actor, task, locale, surface, time, entity availability, or exclusions.

**Alt text**

A ledger compares weak row-count descriptions with a bounded population statement; unresolved checkboxes show that metadata alone cannot verify coverage.

## Slide 04 — Universe → frame → eligible frame → sample → cells

**On-screen text**

Five objects, five denominators:

1. Declared universe
2. Operational frame
3. Eligible frame
4. Selected sample
5. Realized evaluation cells

**Visual specification**

Draw a vertical five-stage sieve. The universe is a wide reservoir. The frame is a narrower aperture with unreachable units drawn outside it. The eligibility stage diverts hatched records to an exclusion ledger. The sample stage selects numbered tokens. The cell stage expands tokens into a matrix. Beside every arrow, place numerator/denominator placeholders and an `unknown` label where a count is unavailable.

**Speaker notes**

Name the loss at each transition. Frame undercoverage occurs when a target unit has no route into the operational source. Eligibility rules remove records for declared reasons. Sampling can leave thin or empty strata. Collection can lose cells to refusal, outage, governance stop, or missing interface state. These are not the same failure. The denominator ledger should preserve planned and realized counts. If a text-only frame lacks voice queries, no later weighting repairs that zero selection path. If a frame contains a rare stratum but the sample misses it, resampling or redesigned allocation may help.

**Teaching check**

Ask where a private-query exclusion belongs and where an unlogged Cantonese voice task belongs. Expected: exclusion between frame and eligible frame; the absent voice task is frame undercoverage.

**Alt text**

A five-stage sieve keeps unreachable, excluded, unselected, and unrealized units visible and labels each transition with its own denominator.

## Slide 05 — Declare the universe in seven dimensions

**On-screen text**

Actor · task · entity · language/locale · surface · time · exclusions

Example target:

> Non-personal English procurement tasks used by Hong Kong small-business technology evaluators during a declared quarter.

**Visual specification**

Create a heptagon with one dimension on each side. Put the example statement in the center. Attach one “not all…” callout to each side: not all consumers, tasks, categories, languages, surfaces, dates, or sensitive requests. Use numbered labels so the long description can follow the same order.

**Speaker notes**

The universe should be broad enough to matter and narrow enough to audit. Actor matters because a procurement officer and casual consumer face different decisions. Task matters because learning and selecting are different intents. Entity scope controls category availability. Language is not interchangeable with locale. Surface matters because API and consumer UI can differ. Time matters for demand and system state. Exclusions protect safety and clarify inference. The example is not a recommended universal universe; it demonstrates explicit boundaries. Learners should resist “all users” unless an actual frame and transport design support it.

**Teaching check**

Ask learners to add one out-of-scope group to the example without changing the target. Expected answers include Cantonese-language users, medical decisions, personal queries, or other markets.

**Alt text**

A seven-sided scope diagram surrounds one bounded population statement; each side names a required dimension and a corresponding out-of-scope population.

## Slide 06 — The frame is an access mechanism

**On-screen text**

Possible frames:

- governed log extract
- survey task generator
- expert-authored matrix
- public benchmark
- synthetic grammar

Each reaches a different population.

**Visual specification**

Show five doors leading into one study room. Each door has a different filter pattern and a small provenance tag. Behind each door, draw populations that can and cannot pass. A footer line reads: “A transparent synthetic frame can be useful without being naturalistic.”

**Speaker notes**

A frame is not automatically a list. It can be a procedure that generates eligible units. A log frame may approximate observed usage but carries governance, missing-channel, and historical-product limitations. A survey generator can target roles but depends on recruitment and task wording. Experts can design diagnostic coverage but overrepresent professionally salient tasks. Benchmarks inherit their source domain and date. Synthetic grammars enable factorial tests but cannot estimate organic frequency. The best frame depends on the estimand. Combining frames requires provenance strata and a declared rule, not simple concatenation.

**Teaching check**

Ask which frame best estimates real query frequency. Correct response: none by label alone; a governed probability-linked usage frame may help, but its channel coverage and selection mechanism must be audited.

**Alt text**

Five differently filtered doors represent five frame mechanisms; each permits only part of a population to enter the study.

## Slide 07 — Sample the intent cluster, not the wording

**On-screen text**

Primary unit: independently interpretable query-intent cluster

One intent may have:

- short phrasing
- constrained phrasing
- branded variant
- validated language variant

Keep the family together when splitting.

**Visual specification**

Place one central intent node above four prompt cards connected by solid family lines. A vertical development/held-out divider runs below; all four cards stay on one side. Show an invalid alternative as a small crossed-out inset with cards split across both sides.

**Speaker notes**

Literal strings are often dependent manifestations of one information need. Treating every paraphrase as independent exaggerates query diversity and can leak the same fact need into development and held-out evaluation. Assign a stable intent-cluster ID before splitting. Controlled variants can remain because wording sensitivity is a legitimate object, but their dependence must be preserved in allocation and uncertainty. A held-out panel is not protected if the development set contains semantic twins. This unit choice also governs later cluster resampling in W07.

**Teaching check**

Ask whether “Compare A and B” and “How does B differ from A?” may be placed in opposite splits. Expected: not if they instantiate the same underlying comparison need.

**Alt text**

Four wording variants connect to one intent cluster and remain on one side of a development–held-out divider; a crossed-out inset shows the invalid split.

## Slide 08 — The sampling cube exposes empty cells

**On-screen text**

Axes:

- intent
- entity class
- locale × time

Selected · thin · empty · out of scope

**Visual specification**

Design an isometric cube composed of labeled cells, paired with a flat table beneath it. Solid dots mark selected cells, open circles thin cells, diagonal hatching empty but in-scope cells, and gray blocks out-of-scope cells. The table repeats every encoding so the cube is nonessential. Do not imitate an external paper figure.

**Speaker notes**

The cube is a planning device, not evidence that all dimensions are fully crossed. High-dimensional designs quickly create sparse cells. Decide which contrasts matter before collection. An empty in-scope cell means unsupported coverage, not zero visibility. An out-of-scope cell is deliberately excluded and should not be invoked in conclusions. Thin cells may support descriptive inspection but not stable subgroup claims. The paired table prevents perspective and color from hiding cell status.

**Teaching check**

Ask how to report an in-scope locale-intent cell with no eligible query. Expected: as an instrument coverage gap; do not impute a zero outcome.

**Alt text**

A three-axis cube and matching table classify cells as selected, thin, empty, or out of scope, making coverage gaps visible.

## Slide 09 — Allocation and weights answer different questions

**On-screen text**

\[
\widehat\mu=\sum_h W_h\widehat\mu_h
\]

- (n_h): sample allocation
- (W_h): target weight
- weight source: design, usage estimate, policy, or equal diagnostic

**Visual specification**

Use four horizontal strata bars. Above them show equal sample tokens: 20 per stratum. Below them show unequal target-weight rulers: 0.50, 0.25, 0.15, 0.10. At right, display two labeled result boxes: “equal-stratum diagnostic” and “declared target-weighted.”

**Speaker notes**

Read the equation as “the overall estimate is the sum of each stratum estimate times its declared target weight.” Oversampling a rare stratum can improve its precision without pretending the stratum is common. Equal allocation is a design choice; equal weighting is an estimand choice. Usage weights require evidence about usage. Policy weights encode decision priority, not prevalence. Record the weight source and freeze it before seeing the outcome. Show both a diagnostic profile and a target-weighted estimate when each is relevant.

**Teaching check**

Ask whether equal numbers of learning and purchasing prompts imply equal demand prevalence. Expected: no; they establish only equal allocation unless independent weight evidence exists.

**Alt text**

Four strata receive equal sample counts but unequal target weights, producing separately labeled diagnostic and target-weighted summaries.

## Slide 10 — Inclusion and exclusion must ignore outcomes

**On-screen text**

Include if:

- in-scope task and locale
- provenance known
- interpretable and non-sensitive
- cluster and version assigned

Exclude by predeclared reason—not by result.

**Visual specification**

Draw a two-lane decision gate. The reviewer sees query metadata and text on the left. A wall hides outcome fields on the right. Decisions flow to “eligible” or a hatched “exclusion log” containing ID, reason code, reviewer, and date. An alarm icon labels any attempt to open the outcome wall.

**Speaker notes**

The reviewer should be able to apply eligibility rules without knowing whether the query mentions, cites, or ranks the target. “Low quality” is too vague; define missing constraint, uninterpretable corruption, out-of-scope task, sensitive data, duplicate intent, unknown provenance, or desired-answer cue. Preserve stable IDs and decisions. If privacy requires text deletion, retain lawful non-sensitive metadata and the governance basis. Never replace a difficult query because it hurts a preferred metric. That would redefine the population around the outcome.

**Teaching check**

Ask why “exclude prompts that receive no citations” is invalid for a citation-rate study. Expected: it conditions eligibility on the outcome and changes the denominator to favorable or citation-bearing cases.

**Alt text**

A blinded eligibility gate separates query information from hidden outcome fields and routes exclusions into a logged record with stable identifiers.

## Slide 11 — Preserve the denominator ledger

**On-screen text**

Planned → framed → eligible → sampled → issued → returned → labeled → analyzed

Every arrow needs:

- count
- reason for loss
- version
- missing-state policy

**Visual specification**

Create an eight-column flow with a main dark line and downward loss branches. Each loss branch ends in a labeled box such as out of frame, policy exclusion, sampling omission, refusal, timeout, abstention, or protocol failure. Use distinct line patterns and numbers rather than a gradient funnel.

**Speaker notes**

Final analyzed rows are conditional on every prior transition. A refusal can be an outcome, an operational error, censoring, or a protocol failure depending on the preregistered rule. It is not automatically zero visibility. Label abstention is different from missing response. Frame exclusion is different from post-collection loss. The ledger makes these populations inspectable. W06 freezes the upstream part; W07 later models repeated observations and missingness. A dashboard count without this flow cannot identify its denominator.

**Teaching check**

Ask whether a timed-out cell may simply be dropped. Expected: only under a declared policy with the first attempt and retry history preserved; otherwise dropping changes the analyzed population invisibly.

**Alt text**

An eight-stage denominator flow sends different losses to labeled branches, keeping frame, collection, labeling, and analysis exclusions distinct.

## Slide 12 — Prompt text is one field in a cell

**On-screen text**

\[
e=\langle q,i,m,s,\ell,g,t,a,c,\omega,r\rangle
\]

Query · intent · system · surface · locale · geography · time · account · conversation · tool/search state · repeat

Unknown stays unknown.

**Visual specification**

Render the tuple as eleven labeled drawers in a specimen cabinet. One drawer contains the prompt. Two contain `unknown` cards. A warning beneath says: “Provider name is not a version lock.” Pair the cabinet with a linear text list for small screens.

**Speaker notes**

Read each field aloud. Two identical prompts are not the same evaluation cell when conversation history, locale, account, surface, or time differs. Some fields will be unavailable on closed platforms. Record `unknown` rather than infer them. Dynamic aliases may change behind a stable label. Session state can alter tool access or personalization. The purpose is not to claim complete control; it is to prevent uncontrolled variation from disappearing inside the word “query.”

**Teaching check**

Ask which field changes when a prompt is issued after ten prior turns. Expected: conversation state, and possibly account or tool state if the surface also changes.

**Alt text**

Eleven labeled drawers form an evaluation-cell record; the prompt occupies only one drawer and unavailable state fields are explicitly marked unknown.

## Slide 13 — Translation is not a locale experiment

**On-screen text**

Translation may change:

- meaning and politeness
- entity availability
- cultural assumption
- surface and retrieval access

Analyze added locales separately before pooling.

**Visual specification**

Show two parallel paths beginning from one task card. Path A goes through English/Hong Kong review to Surface A. Path B goes through translation, local review, another account and Surface B. Six changed-variable tags hang between paths. A bracket labels the raw comparison “compound contrast.”

**Speaker notes**

Language and locale should have separate fields. Literal translation may fail task equivalence; transcreation may preserve intent while changing words. A different surface, geography, account, or date creates additional contrasts. A single observed answer difference cannot isolate language. A multilingual extension needs a frame statement, qualified local review, equivalence rubric, entity-availability check, and separate reporting. Pooling is a later analytic decision requiring comparable measurement and target weights.

**Teaching check**

Ask for three variables changed when an English web-UI query in Hong Kong is translated and issued through a Chinese API endpoint next month. Accept language, locale, geography, surface, account, time, model state, or tool access.

**Alt text**

Two query paths differ in translation, review, surface, account, geography, and time, so their answer contrast cannot identify language alone.

## Slide 14 — Time and session create distinct instruments

**On-screen text**

Record separately:

- query content-validity window
- system observation time
- account and history
- search/tool state
- retry and repetition

**Visual specification**

Use two synchronized timelines. The upper line tracks the query's factual validity and wording versions. The lower line tracks system observations, session resets, and tool-state changes. Vertical connectors show which query version enters which cell. Broken connectors mark unknown state.

**Speaker notes**

“Current policy” can become stale even if the string does not change. Preserve the old query version and define its validity interval. System time is another variable: retrieval indexes, model aliases, interfaces, and policies can drift. Session state may preserve previous turns or personalization. Retries do not erase the first attempt. These fields prepare W07's repeated-observation analysis. W06 ensures that a repeat means what the design says it means.

**Teaching check**

Ask whether rewriting an August query in place to reflect a September product name is acceptable. Expected: no; create a new version and retain the prior text and mapping.

**Alt text**

Parallel query-validity and system-observation timelines link versioned prompts to session-specific cells while leaving unknown states visible.

## Slide 15 — Provenance is part of the stratum

**On-screen text**

Naturalistic · transformed · expert-authored · template-expanded · translated · model-generated · control

Fluent ≠ naturalistic  
Balanced ≠ prevalent

**Visual specification**

Create seven provenance lanes flowing into one bank. Each lane carries a distinct border pattern and required metadata tag. Before the combined bank, place a “retain provenance” gate. On the far right, separate reporting panels remain visible rather than collapsing into one total.

**Speaker notes**

Every query should carry a provenance class and construction record. Naturalistic data need lawful source, date, inclusion mechanism, and transformation history. Model-generated data need generation instruction, identifiable system state where available, selection procedure, and human review. Expert authorship can produce excellent diagnostic coverage while missing ordinary user phrasing. Synthetic prompts are useful for stress tests and teaching, but they do not estimate demand frequency. Combine provenance sources only under a stated sampling and weighting rule.

**Teaching check**

Ask whether a model-generated query can become naturalistic after a human copy edit. Expected: no; it remains synthetic or transformed synthetic, with the edit recorded.

**Alt text**

Seven provenance lanes enter a query bank through a gate that preserves their labels, and separate analysis panels prevent silent pooling.

## Slide 16 — Duplicate detection surfaces candidates

**On-screen text**

\[
J(A,B)=\frac{|A\cap B|}{|A\cup B|}
\]

Exact normalization + token-set Jaccard + human adjudication

Threshold is a declared convention, not semantic truth.

**Visual specification**

Display two token circles with shared words in the intersection and unique words outside. Beneath, show a three-step track: normalize, score, adjudicate. Place two error callouts: “paraphrase missed” and “controlled variant falsely flagged.”

**Speaker notes**

Explain the score in words: shared unique tokens divided by all unique tokens. The L02 English-oriented rule is transparent and low-compute. It cannot determine semantic equivalence. It may miss low-overlap paraphrases and flag legitimate controlled variants. The report should preserve pair IDs, method, threshold, score, and reviewer decision. Cluster identity is more important than automatic deletion. Another language or tokenizer needs its own validation.

**Teaching check**

Ask whether a pair above 0.80 must always be deleted. Expected: no; it becomes a review candidate, and the codebook determines whether it is a duplicate or intentional variant.

**Alt text**

Two token sets illustrate Jaccard overlap, followed by normalization, scoring, and human adjudication with both false-negative and false-positive risks.

## Slide 17 — Split by intent before tuning

**On-screen text**

Development → validation → held-out

- split intent families, not strings
- freeze before optimization
- one held-out look
- reuse converts test to development

**Visual specification**

Use three locked compartments. Intent-family bundles are allocated wholly to one compartment. A one-way audit hatch opens from held-out to final report. A red return arrow from held-out to editing is labeled “panel consumed.”

**Speaker notes**

The development panel supports prompt, system, and intervention decisions. Validation supports selection among candidates under a fixed rule. The held-out panel estimates performance after choices are frozen. If its result triggers another edit, those queries have entered development and cannot remain the sole final test. Splitting surface paraphrases across compartments leaks the same information need. Hashes and access records support the boundary; they do not substitute for role discipline.

**Teaching check**

Ask what happens when a poor held-out result leads to one more rewrite. Expected: the held-out panel is consumed as development evidence; a new independent test route is needed for a clean final claim.

**Alt text**

Intent-family bundles occupy separate development, validation, and held-out compartments; a return arrow from held-out to editing marks the test as consumed.

## Slide 18 — Leakage has several routes

**On-screen text**

Desired answer · treatment · split · test reuse · label · temporal · provenance

Leakage is information crossing a protected design boundary.

**Visual specification**

Draw seven pipes crossing three walls: query construction, evaluation, and labeling. Each pipe has its leakage name and a shutoff control: blind, freeze, cluster, restrict, timestamp, or label provenance. A small detector icon makes clear that keyword scanning catches only one subset.

**Speaker notes**

Desired-answer leakage places the preferred result inside the query. Treatment leakage makes evaluation wording match the changed source. Split leakage shares intent families. Test reuse allows outcomes to guide another decision. Label leakage reveals rank, system, treatment, or expected direction to judges. Temporal leakage introduces future information into a historical task. Provenance leakage presents synthetic construction as observed demand. Controls are procedural: freeze artifacts, separate access, blind assessment, cluster splits, timestamp records, and audit provenance. Keyword flags are helpful for obvious cues but cannot detect all presuppositions or loaded comparisons.

**Teaching check**

Read “Which provider besides Acme is worth considering?” Ask which leakage risk applies if unprompted Acme mention is the outcome. Expected: desired-answer leakage because Acme is supplied in the premise.

**Alt text**

Seven labeled leakage pipes cross protected query, evaluation, and labeling walls, each paired with a procedural shutoff control.

## Slide 19 — Qrels are judged relations, not ground truth everywhere

**On-screen text**

Qrel record:

topic × object version × rubric × grade × assessor × date

Not:

- query frequency
- universal relevance
- platform preference
- user satisfaction

**Visual specification**

Place a query-topic card and versioned document card on opposite sides of a judgment bridge. The bridge includes rubric, assessor, date, and grade. Four arrows pointing outward are blocked and labeled with the invalid interpretations above.

**Speaker notes**

A qrel is a protocol record. Relevance can depend on task, locale, time, passage granularity, and materiality rule. Record the scale, assessor instructions, blinding, adjudication, abstention, and treatment of unjudged pairs. In a closed synthetic fixture, unlisted pairs may be treated as gain zero if declared. In incomplete pooling, unjudged is not automatically nonrelevant. Qrels do not validate how queries were sampled; the demand instrument and relevance instrument are distinct.

**Teaching check**

Ask whether a grade-2 qrel proves a document should rank highly for every user. Expected: no; it records a bounded judgment under one topic, object version, rubric, assessor process, and date.

**Alt text**

A judgment bridge connects one query topic to one versioned object through rubric, assessor, date, and grade, while blocking four broader interpretations.

## Slide 20 — Govern labels before seeing rankings

**On-screen text**

Freeze:

- codebook and materiality
- judge access
- abstention
- disagreement and adjudication
- unjudged policy
- development vs held-out labels

**Visual specification**

Create a two-person adjudication workflow. Two blinded assessors label independently, disagreements enter a separate diamond, and an adjudicator resolves or preserves abstention. Ranked output is shown behind an opaque curtain until the label freeze. A version stamp seals the final qrel set.

**Speaker notes**

Labels include qrels, intent, entity class, risk, and leakage status. Each is a constructed measurement needing a codebook. When relevance judges see system identity, rank order, or treatment, expectations can influence labels. Blind where feasible and record when blinding is impossible. Agreement measures reproducibility under the rubric, not truth or validity. Preserve disagreement and abstention. Keep development qrels separate from held-out qrels, and disclose pooling systems if retrieval outputs contribute candidates.

**Teaching check**

Ask why judges should not revise a qrel after seeing that a preferred retriever misses the document. Expected: it contaminates the evaluation label with method outcome and can favor the preferred system.

**Alt text**

Two blinded assessors feed disagreements to adjudication while ranked outputs stay hidden until a versioned qrel set is frozen.

## Slide 21 — L02: 20 inputs become 18 accepted records

**On-screen text**

FRAME-DEMO-001 · synthetic · en-HK · threshold 0.80

| Input | Accepted | Excluded | Duplicate pairs |
|---:|---:|---:|---:|
| 20 | 18 | 2 | 1 |

Q-019 near duplicate of Q-004: 0.818182  
Q-020: desired-answer leakage

**Visual specification**

Use a denominator waterfall with 20 numbered tokens. One token moves to a near-duplicate tray, one to a leakage tray, and 18 remain in six labeled intent bins. Under the diagram print counts: development 13, held-out 5, locale en-HK 18. A warning banner says “fewer than 60 required for full assignment.”

**Speaker notes**

These are reproduced fixture facts. The code passes structurally because no contract error occurs, while the report retains a scientific warning about size. Q-004 and Q-019 have token-set Jaccard 0.818182 under the declared tokenizer; the fixture excludes the later record. Q-020 explicitly prescribes the desired ranking and is excluded for leakage. The accepted set contains learn 3, compare 3, select 3, verify 4, counterfactual 3, and negative-control 2. The result demonstrates the audit workflow. It does not estimate real query frequency or validate 0.80 universally.

**Teaching check**

Ask why `PASS` and a warning can coexist. Expected: the software contract is valid, while scientific readiness for the 60-query full assignment is not yet met.

**Alt text**

A waterfall retains 18 of 20 synthetic queries, diverting one near duplicate and one desired-answer prompt; intent, split, locale, and size-warning counts remain visible.

## Slide 22 — Read evidence cases through their sampling boundary

**On-screen text**

- **PAPER-29:** bounded repeat panel; no universal run count
- **PAPER-38:** generated prompts and changing denominators; no organic-demand claim
- **PAPER-26:** fixed-context variants; no live query-planning claim
- **PAPER-01:** time-provenance quarantine
- **PLAT-04:** one preview interface; no cross-platform sampling standard

**Visual specification**

Create five evidence cards in a single row. Each card has three lines: useful design feature, population boundary, forbidden transfer. Put PAPER-01 behind a diagonal quarantine pattern. Put PLAT-04 in a distinct official-document outline. Do not reproduce any source title page or chart.

**Speaker notes**

The readings are cases to audit, not templates to copy. PAPER-29 motivates repeated measurement but its prompt panel, language, markets, systems, and finite run analysis bound its precision results. PAPER-38 provides a large vendor monitoring case while using generated prompts and multiple changing analytic universes. PAPER-26 uses generated variants in a supplied fixed context and does not observe live demand. PAPER-01 has unresolved dates beyond the course freeze and remains quarantined. PLAT-04 documents named fields on one Microsoft preview. Each source helps identify a design choice only inside its evidence ceiling.

**Teaching check**

Ask which source can establish the prevalence of procurement queries among all users. Expected: none of these sources, absent a suitable population sampling design.

**Alt text**

Five bounded evidence cards pair a legitimate teaching use with a population limit; PAPER-01 is visibly quarantined and PLAT-04 is marked as official interface documentation.

## Slide 23 — Two outcome-equivalent learning routes

**On-screen text**

Seminar: 90 planned minutes  
Studio: 75 planned minutes  
Async core: approximately 210–220 planned minutes  
Transcript equivalent: 35–45 planned minutes

No W06 recording or pilot exists.

**Visual specification**

Show parallel synchronous and asynchronous lanes converging on the same seven artifacts. The synchronous lane contains seminar and studio blocks. The asynchronous lane contains notes, case, L02 run, and transcript blocks. Use planning-budget labels rather than clock timestamps. Place a clear status bar: “recording: none; pilot: pending.”

**Speaker notes**

Accessibility means equivalent scientific content and decisions, not identical media. The transcript contains every core proposition needed without video. Slide visuals have alt text and table equivalents. The L02 route uses the Python standard library and can also be inspected from supplied outputs. All durations are authoring budgets. They have not been observed in a class, rehearsal, recording, or usability study. A future pilot must record actual timing, questions, errors, accessibility findings, and grading behavior before the route can be described as validated.

**Teaching check**

Ask whether the transcript may be called a recording transcript. Expected: no; it is a planned readable/recordable equivalent until real synchronized media exists.

**Alt text**

Synchronous and asynchronous learning lanes converge on identical artifacts; a status bar states that timings are planned and no recording or pilot exists.

## Slide 24 — Exit ticket: state coverage and absence

**On-screen text**

Complete both sentences:

1. “This instrument represents ___ under ___ because ___.”
2. “It does not represent ___, and the strongest unresolved frame gap is ___.”

Freeze before answer collection.

**Visual specification**

Return to the lens from Slide 01. This time label the reached region with actor, task, locale, time, and provenance. Label three unreached regions explicitly. Place a small checksum seal on the prompt stack and an arrow forward to W07: “estimate only after scope is frozen.”

**Speaker notes**

End with a bounded conclusion rather than a metric. For the demonstration fixture, a valid statement is that 18 of 20 synthetic English teaching queries were accepted under FRAME-DEMO-001 after one declared near-duplicate decision and one desired-answer exclusion. It is invalid to call those 18 representative of real demand. Each learner must name represented scope, conditions, evidence for entry, unrepresented populations, and the largest uncertainty. The honest output may be “instrument not ready.” That is a successful diagnosis when the frame cannot support the intended inference.

**Teaching check**

Collect the two sentences. Reject any answer that omits the frame source, locale/time condition, or at least one unrepresented population.

**Alt text**

A labeled sampling lens reaches a bounded portion of the task population, leaves three regions explicitly unreached, and sends a checksum-sealed instrument forward to later estimation.
