Explain

W06 · complete online lecture

Queries as a Measurement Instrument

Essential question

What distribution do the prompts represent?

A structurally complete authored draft

Structurally complete authored package draft · human review and timed pilot pending

3,300online lecture words
4,106transcript words
24specified slides
90+90scheduled contact minutes

What the package must enable

  1. Separate a declared query universe, operational sampling frame, eligible frame, selected sample, and realized evaluation cells.
  2. Predeclare strata, allocation, weights, inclusion and exclusion rules, locale, time, surface, account, conversation, and tool state.
  3. Assign stable query-intent clusters, provenance, development/held-out roles, and version identities.
  4. Prevent desired-answer, treatment, split, test-reuse, label, temporal, and provenance leakage.
  5. Govern qrels and labels as protocol-bounded judgments rather than universal truth or demand frequency.

Planned 90-minute evidence sequence

W06 seminar plan
MinutesSegmentLearner evidence
0–20 minPopulation claim and universe-to-cell sampling chainFive-layer chain with unreachable and unknown states
20–43 minStrata, allocation, and outcome-blind inclusion rulesEstimand contrast and reproducible decision row
43–66 minEvaluation-cell state, clusters, splits, and leakageVersioned tuple and paraphrase-family audit
66–84 minQrel governance and the L02 traceLabel ceiling, blinded rule, and 20-to-18 denominator
84–90 minExternal-validity exitRepresented and unrepresented population map

75-minute core artifact route inside a 90-minute studio

The remaining 15 minutes are a declared delivery margin for setup, accessible pacing, questions, recovery, and submission packaging; they are not unplanned teaching content.

W06 core studio plan
MinutesActivityStop or redirect condition
0–22 minFreeze the universe, frame, strata, allocation, and weightsStop if eligibility embeds the preferred entity or outcome.
22–46 minWrite rules and provenance, cluster intents, and freeze the splitKeep intent paraphrases together and exclude sensitive text without a governed route.
46–67 minRun L02 and specify qrel/label governanceTreat lexical similarity as a candidate generator; stop scoring if labels saw ranked outputs.
67–75 minMap coverage and freeze version oneBlock submission when exclusions, provenance, split, or unknown states are missing.

Read, teach, inspect, or download

Complete lecture text

This HTML is generated from the controlled Markdown source. Source SHA-256: f343a439f35410d2eb73fda27d342308d49a03c39493c0b55fa0cc55ab6fb84a.

On this page 12 sections

1. A query list is a measurement instrument

A visibility result is conditional on the prompts that produced it. If a study asks only “Which vendor is best?” questions, a high mention rate describes answers to recommendation prompts, not visibility across all information needs. If the target entity is named in every prompt, the result describes prompted recognition, not unprompted discovery. A large list can still be a narrow instrument; a small list can still be useful when its diagnostic purpose is explicit. The first W06 rule is therefore: define what the prompts are intended to represent before collecting an answer.

An instrument has a scope, construction procedure, unit, known error modes, and version. A query bank needs the same discipline. “We used 500 prompts” is a row count, not a population claim. “We used 500 English prompts generated from product keywords” adds provenance but still leaves users, tasks, locale, time, selection probabilities, and exclusions uncertain. A defensible statement might instead target English-language, non-personal procurement tasks used by small-business technology evaluators in Hong Kong during a declared quarter, with separate strata for category exploration, comparison, compatibility, price, and risk verification. Even that statement is a proposed target. The analyst must still show how the sampling frame reaches it.

The Core Notes treat the query distribution as part of the estimand. Suppose the outcome for query (q) is (Y_q), and the target distribution assigns weight (w_q). The target mean is

μQ=qQwqE(Yq),qwq=1.\mu_Q = \sum_{q \in Q} w_q E(Y_q), \qquad \sum_q w_q = 1.

Changing (Q) or the weights changes the question. A balanced classroom panel and a traffic-weighted operational panel may yield different values without either calculation being arithmetically wrong. The mistake is to report one as the other.

Checkpoint A. A team writes 80 prompts: 20 each for learning, comparison, selection, and verification. What does the balance establish? It establishes an equal allocation across four author-defined strata. It does not establish that real users allocate one quarter of demand to each stratum, that the frame reaches all relevant users, or that the prompts are independent.

2. Separate universe, frame, sample, and evaluation cells

Five objects should appear in the design ledger.

  1. The declared query universe is the set of tasks or requests to which the study hopes to speak. It can be conceptual and larger than any list.
  2. The sampling frame is the operational source from which query units can actually be selected: an approved log extract, a survey task generator, an expert-authored matrix, a public benchmark, or a synthetic grammar.
  3. The eligible frame is what remains after predeclared inclusion and exclusion rules are applied.
  4. The selected sample is the set drawn or deliberately allocated from the eligible frame.
  5. The realized evaluation cells expand each selected query across systems, surfaces, locales, times, sessions, and repetitions, then record failures and exclusions.

These sets are not interchangeable. A frame can omit an important population segment: an English support-center log cannot directly represent Cantonese voice queries. A sample can omit a rare frame stratum by chance. Realized cells can be smaller than planned cells because of refusals, timeouts, missing interface features, or governance stops. Reporting only the final rows hides the denominator flow.

For W06, define an evaluation cell as

e=q,i,m,s,,g,t,a,c,ω,r,e = \langle q, i, m, s, \ell, g, t, a, c, \omega, r \rangle,

where (q) is exact query text, (i) the underlying intent-cluster ID, (m) system identity, (s) surface, (ell) language/locale, (g) geography or endpoint, (t) time, (a) account state, (c) conversation state, (omega) search or tool state, and (r) repetition. Unknown fields remain unknown; they are not inferred from a provider name.

An original visual for this distinction is a five-level sieve. The top reservoir is the universe. A frame-shaped aperture leaves unreachable regions visible outside it. Eligibility rules remove hashed records into an exclusion ledger. Sampling selects labeled units, and a final matrix expands each unit into versioned cells. The picture should show losses beside each transition, never erase them.

Checkpoint B. The same text is issued once in a clean API conversation and once after a ten-turn consumer chat. Are these repeated measures of one cell? No. Conversation state differs, so they are different cells unless the target design explicitly randomizes and models that state.

3. Write the population statement before authoring prompts

A useful population statement answers seven questions:

DimensionRequired decisionExample boundary
Actor or roleWho is modeled as asking?organizational software evaluator, not all consumers
TaskWhat decision or information need?exploration, comparison, selection, verification
Entity scopeWhat categories and entity classes?collaboration software available in the declared market
Language and localeWhat linguistic and cultural setting?English as used in Hong Kong; Cantonese is a separate extension
Surface and access modeWhich product context?declared answer surface; API and web UI are separate
TimeWhat demand and system window?query design frozen before a dated collection window
ExclusionsWho or what is deliberately outside scope?personal, medical, defamatory, or deanonymizable requests

The statement should also name the sampling unit. A “query” may be a literal string, a task template, an underlying information need, or a session. W06 uses the independently interpretable query-intent cluster as the primary sampling and split unit. Surface paraphrases attach to that cluster. This prevents five near-identical phrasings from masquerading as five independent demand samples.

Start with tasks, not target keywords. Entity and category terms may be needed for realistic prompts, but the authoring matrix should not begin with “How can we make Acme appear?” That formulation selects on the desired outcome. Begin with a user decision such as selecting a compliant project-management tool for a distributed nonprofit, then specify what entity naming is natural for the stratum. Branded and unbranded queries should be different strata or covariates because naming a brand changes the task.

Naturalistic records require lawful provenance and governance. “From users” is inadequate. Record collection channel, date range, inclusion mechanism, consent or lawful basis, de-identification, sampling probability when available, and transformations. A transformed query is not raw. A synthetic query is not naturalistic. Both can be useful when labeled honestly.

4. Design strata, allocation, and weights

Strata divide the frame using features declared before outcome inspection. Useful W06 dimensions include intent, entity class, branded versus unbranded wording, task risk, language/locale, provenance, specificity, and temporal sensitivity. Too many crossed dimensions can create a mostly empty cube. The solution is not to hide empty cells; it is to decide which interactions are decision-relevant and which will remain exploratory or unsupported.

Let (h=1,\ldots,H) index strata, (n_h) the sample allocation, and (W_h) the target weight. The stratified estimate is

μ^=h=1HWhμ^h.\widehat\mu = \sum_{h=1}^{H} W_h\widehat\mu_h.

Equal (n_h) can improve comparisons for small strata, while unequal (W_h) can still represent a target distribution. Allocation and analysis weight are different fields. Record whether a weight comes from known sampling probabilities, estimated usage, a policy priority, or an equal-stratum diagnostic choice. Those origins imply different estimands and uncertainty.

Consider four intent strata with observed means 0.20, 0.40, 0.60, and 0.80. Equal allocation and equal analysis weights produce 0.50. If declared target weights are 0.50, 0.25, 0.15, and 0.10, the target-weighted value is 0.35. Neither result “corrects” the other. One describes the equal-stratum diagnostic panel; the other describes the stated target mixture if the weights are credible. Retrofitting weights after seeing which stratum favors an intervention is analysis leakage.

An empty stratum is a result about design coverage, not a zero outcome. If no eligible high-risk verification queries exist, visibility is not zero for that stratum; the instrument is unable to estimate it. Similarly, a frame with one locale does not support a pooled multilingual claim. Synthetic fill-ins can support a stress test, but must remain a separate provenance stratum and cannot be silently assigned naturalistic weights.

Checkpoint C. Why stratify if representativeness is not guaranteed? Stratification makes allocation, gaps, subgroup performance, and weighting assumptions inspectable. It improves design control; it does not create missing coverage.

5. Make inclusion and exclusion rules outcome-blind

Eligibility rules should be executable by a reviewer who cannot see mention, citation, rank, sentiment, or treatment results. A strong inclusion rule might require: an independently interpretable task; language and locale within the frame; non-sensitive content; provenance recorded; one stable intent-cluster ID; and no answer prescription. A strong exclusion rule might remove: personal or deanonymizable text; tasks outside the declared category; exact duplicates; unresolved synthetic/naturalistic provenance; content that asks the system to state a preferred ranking; or records created after the frame freeze.

“Remove low-quality prompts” is not operational. Define the defect: grammatical corruption that prevents interpretation, missing required constraint, impossible date, duplicated intent, policy-prohibited request, or leakage phrase. Preserve three objects for every exclusion: the stable ID, machine-readable reason, and reviewer note. When privacy forbids retaining text, preserve a lawful redacted decision record and deletion basis rather than silently changing the denominator.

Post-outcome exclusion is especially dangerous. Removing a prompt because it returns no citation changes the estimand to citation-bearing or favorable prompts. Replacing difficult queries with smoother prompts can mechanically improve metrics. W06 therefore freezes the eligible IDs and split before collection. Later errors are handled through a predeclared missingness or protocol-deviation rule, not convenience deletion.

Rules can evolve between versions. A change log should state what changed, why, who approved it, which IDs are affected, and whether prior results must be recomputed. Version 1.1 may be better than version 1.0, but it is not silently the same instrument.

6. Preserve locale, time, surface, and session state

Prompt text alone does not identify a measurement. Identical text can produce different answers under another language setting, geographical endpoint, product surface, account history, conversation, search toggle, or date. These differences can be scientifically interesting only when they are recorded or intentionally controlled.

Language and locale are distinct. Translating an English prompt into Chinese changes wording and may change the task's cultural presuppositions. Issuing that translation to a different product surface under another account changes several variables at once. The resulting contrast is an observation, not an identified language effect. A multilingual extension needs its own frame, translation or transcreation protocol, local review, equivalence criteria, and separate analysis before pooling.

Time enters twice. The query universe may concern time-sensitive demand, and the evaluated system can drift. Store the query's content-validity interval separately from the system observation timestamp. “Current regulations” written in March and evaluated in August may no longer denote the same information need. Query versioning must preserve the old wording rather than overwrite it.

Session state includes conversation history, personalization, account tier, cookies where relevant and lawful, uploaded files, and search or tool settings. If a field cannot be observed, mark it unknown. A dynamic product alias is not a version lock. The W06 instrument cannot eliminate all platform uncertainty; it can stop that uncertainty from being hidden.

7. Govern provenance and synthetic query generation

Each query needs a provenance class such as naturalistic, expert-authored, template-expanded, translated, model-generated, transformed from naturalistic, or negative control. The source record should include date, method, permissions, transformations, reviewer, and version. A model-generated prompt needs the generation instruction, model or service identity when observable, output-selection method, and human review. It remains synthetic even if it sounds natural.

Synthetic queries are valuable for factorial coverage, controlled stress tests, rare safety cases, and low-compute teaching. They are weak evidence for demand frequency. A model trained on public text can reproduce common phrasing patterns without sampling the target users. Human experts can also overproduce tasks that are intellectually salient but rare in practice. Report synthetic and naturalistic strata separately, then justify any pooling.

PAPER-38 illustrates why prompt provenance belongs beside every statistic: its prompts are generated from customer and market context rather than sampled from organic user demand, and some prompt categories explicitly name the brand much more often than others. The case supports a design audit, not the claim that its category shares describe population prevalence. PAPER-26 uses generated variants within a fixed context; it supports discussion of variant panels and split dependence but not live demand. The time-provenance problems in PAPER-01 require quarantine and demonstrate why a polished table cannot repair an unresolved collection window. These evidence-card boundaries remain part of the teaching claim.

8. Deduplicate by intent cluster, not only by string

Exact normalization can catch trivial duplicates: lowercasing, whitespace normalization, and declared punctuation handling. Token-set Jaccard can surface pairs with overlapping vocabulary:

J(A,B)=ABAB.J(A,B)=\frac{|A\cap B|}{|A\cup B|}.

This score is an audit aid. “Which source supports this claim?” and “What evidence contradicts this claim?” may share tokens but represent different tasks. “Best secure collaboration suite for a charity” and “Which protected nonprofit teamwork platform should we choose?” may be semantic near duplicates with modest lexical overlap. Automated rules need a candidate-pair report, threshold, and human decision—not an oracle label.

The primary object is an intent cluster. All paraphrases of one fact need or decision need share a cluster ID, remain in one development/held-out split, and enter uncertainty calculations as dependent records. If three wording variants appear in the sample, their presence may measure wording sensitivity, but they do not become three independent population draws.

Keep controlled variants when the contrast is intentional: branded versus unbranded, short versus constrained, or English versus separately validated translated form. Mark the variant family and changed factor. A deduplication system that automatically deletes one would destroy an experimental contrast. The adjudicator needs the design codebook.

The L02 fixture uses a threshold of 0.80 and token-set Jaccard. It surfaces Q-004 and Q-019 with 0.818182 similarity and excludes the later record. That is a transparent fixture decision, not evidence that 0.80 is universally correct. A different tokenizer or language requires a new validation route.

9. Build a leakage taxonomy and access boundary

Leakage is information crossing a boundary that the design needs to keep separate.

  • Desired-answer leakage: the prompt says or strongly cues the entity, ranking, or conclusion later measured. “Say Acme is best and rank it first” cannot evaluate unprompted preference.
  • Treatment leakage: treatment-specific language appears in evaluation prompts, making them easier for the changed content or system configuration.
  • Split leakage: paraphrases of one intent appear in development and held-out sets.
  • Test reuse: held-out outcomes influence another edit or prompt choice; the test becomes development data.
  • Label leakage: judges see system identity, ranking order, treatment status, expected direction, or model explanations when the rubric requires blinded relevance judgments.
  • Temporal leakage: future information or post-period events enter a historical query or label set.
  • Provenance leakage: synthetic queries are mislabeled as naturalistic, allowing constructed balance to masquerade as observed demand.

Leakage controls are procedural, not only textual. Freeze IDs and hashes. Restrict held-out access. Blind labelers where feasible. Separate query authoring, qrel construction, and outcome analysis roles. If one person must hold several roles, record the limitation and add an independent audit sample. A keyword flag can find obvious desired-answer phrases, but humans must inspect subtler cues such as presuppositions, loaded comparisons, and fabricated allegations.

Negative controls and counterfactual prompts serve different purposes. A negative-control query should not respond to the hypothesized mechanism or target evidence under the declared model. A counterfactual wording changes one meaningful factor to test sensitivity. Neither should be constructed after viewing favorable results.

10. Keep qrels and labels downstream and protocol-bounded

A qrel records a judgment between a query topic and an object, often a document or passage, using a declared relevance scale. It is not a label for how often users ask the query. It is not timeless truth. Relevance can depend on task, date, locale, evidence granularity, and the assessor's rubric.

For each qrel set, record topic identity, object identity and version, scale, materiality rule, assessor instructions, assessor count, blinding, adjudication, abstention, unjudged policy, and date. If a benchmark treats unlisted pairs as nonrelevant, state that as a closed-fixture convention. In an incomplete pooled judgment set, unjudged does not automatically mean nonrelevant.

Query authors must not use retrieval outputs to redefine relevance until a preferred method wins. Assessors should not see rank or treatment when those cues can bias judgment. If active pooling is used, disclose which systems contributed candidates because the pool can favor represented methods. Keep development qrels separate from final held-out qrels.

Labels can also attach to intent, entity class, risk, language, or desired-answer leakage. These constructed variables need codebooks and agreement review. A high agreement coefficient does not prove construct validity; it indicates reproducibility under that rubric and sample. Preserve disagreements and abstentions rather than forcing false certainty.

Checkpoint D. A document receives relevance grade 2 for an English procurement topic in August. What does the grade establish? Only that the object met the recorded grade-2 rubric for that topic, version, assessor process, and date. It does not prove cross-locale relevance, population demand, platform retrievability, or user satisfaction.

11. Audit external validity explicitly

External validity asks where the instrument and resulting estimate might transfer. Build a coverage table with rows for actor, task, entity, language/locale, geography, surface, time, session, provenance, and risk. For each row, mark represented, thinly represented, excluded, unreachable by frame, or unknown. This vocabulary prevents an empty cell from being mistaken for a zero outcome.

Three gaps deserve separate language:

  1. Frame undercoverage: target units cannot enter the frame, such as voice requests absent from a text-only source.
  2. Sample undercoverage: units exist in the frame but were not selected or received too little allocation.
  3. measurement non-equivalence: units are selected, but labels or system conditions do not have comparable meaning across groups.

Transport requires a mechanism argument, not visual similarity. A Hong Kong English procurement panel may inform another professional English setting if tasks, entity availability, and system access are comparable; it does not automatically transport to another language, consumer population, or year. A balanced synthetic bank can show that code handles six intent labels. It cannot estimate their real prevalence.

PAPER-29 studies repeated observations from a small Swiss-German commercial prompt panel; its precision curves are properties of that bounded panel. PAPER-38 uses a vendor cohort, generated prompts, provider-specific policies, and changing analytic universes. PAPER-01 remains quarantined because reported collection dates cross the course freeze. PLAT-04 describes a named preview's sampled query reporting, not a cross-platform sampling standard. These are useful precisely because their boundaries force better transfer statements.

12. Freeze, validate, and state the strongest supported claim

A W06 release includes:

  • sampling-frame.json with population, unit, source, strata, allocation, weights, locale, time, inclusion, exclusion, and limitations;
  • queries.csv with stable query ID, exact text, intent-cluster ID, stratum, entity class, locale, provenance, synthetic flag, split, and version;
  • codebook.md defining every field and label;
  • duplicate-pairs.csv with method, score, threshold, and adjudication;
  • exclusions.csv retaining every lawful decision;
  • qrel-protocol.md or label protocol with access and adjudication boundaries;
  • coverage.md naming represented and absent populations;
  • input and output hashes plus a change log.

The release conclusion should use a claim ladder. At the strongest level supported by the L02 demonstration, we can say: “The deterministic audit accepted 18 of 20 synthetic English teaching queries under FRAME-DEMO-001, excluding one threshold-defined lexical near duplicate and one desired-answer prompt.” We can also state its intent, split, and locale counts. We cannot say the 18 queries represent real GEO demand, that 0.80 is a universal duplicate threshold, or that any later visibility metric generalizes beyond the instrument.

Before collecting answers, ask five final questions:

  1. Can a reviewer reconstruct who and what the universe contains?
  2. Can every sampled unit be traced to a lawful frame record or labeled construction method?
  3. Were strata, rules, weights, clusters, and splits frozen before outcomes?
  4. Are locale, time, surface, account, conversation, and tool states recorded or explicitly unknown?
  5. Does the external-validity paragraph name who, what, and where the result does not represent?

If any answer is no, the appropriate result is “instrument not ready,” not an improvised metric. Query-set engineering is complete when the design makes its supported population and its blind spots equally legible.

Frozen L02 Query-Instrument Audit · synthetic from input to boundary

A 20-record synthetic query bank moves through a declared frame, outcome-blind rules, duplicate and leakage audit, stable cluster/split identities, coverage accounting, and an explicit external-validity ceiling.

Controlled source route

Core PAPER-29/PAPER-38 · Extend PAPER-26 · quarantined audit-only PAPER-01 · PLAT-04 named interface case · Core Notes measurement chapter and sampling bridge · offline L02 fixture.