# W06 Transcript Equivalent — Queries as a Measurement Instrument

## Status and use

No W06 recording exists. This document is a planned readable or recordable equivalent for learners who cannot or do not use video. The chapter durations below are **instructional planning budgets**, not timestamps recovered from media and not observed speaking times. Every core scientific proposition needed for the W06 outcomes appears in this text. A future recording would require separately verified timing, captions, transcript synchronization, audio description, and accessibility review.

## Planned chapter budget

| Chapter | Planned duration | Focus |
|---|---:|---|
| 1 | 4 minutes | Why prompts are an instrument |
| 2 | 4 minutes | Universe, frame, sample, and cells |
| 3 | 4 minutes | Population statements and sampling units |
| 4 | 4 minutes | Strata, allocation, and weights |
| 5 | 4 minutes | Inclusion, exclusion, and denominator flow |
| 6 | 4 minutes | Locale, time, surface, and session state |
| 7 | 4 minutes | Provenance, duplicates, and intent clusters |
| 8 | 4 minutes | Leakage and split discipline |
| 9 | 4 minutes | Qrels and label governance |
| 10 | 4 minutes | L02 worked audit and evidence routes |
| 11 | 2 minutes | External-validity exit and handoff |

**Total planned route: 42 minutes.** This total falls inside the intended 35–45 minute no-video-equivalent range, but remains unpiloted.

## Chapter 1 — Why prompts are an instrument

Imagine that a report begins with the sentence, “We tested five hundred prompts.” What have we learned? We know a nominal row count. We do not yet know who those prompts represent, what tasks they express, how they were selected, which languages and locales they cover, whether they name the target, or what kinds of requests could never enter the list. Five hundred is large compared with five, but it is not a population definition.

A visibility estimate is conditional on its prompts. If all prompts ask for the best provider, the estimate describes recommendation-style tasks. If all prompts name one brand, a mention rate describes prompted recognition rather than unprompted discovery. If a team generated prompts from its own product keywords, the instrument may overrepresent the concepts that the team already publishes. No arithmetic applied later can erase those conditions.

That is why W06 treats the query set as a measurement instrument. An instrument needs a declared target, an operational construction process, a unit, a version, and known error modes. A thermometer has a range and calibration; a query bank has a universe, frame, inclusion rule, strata, provenance, split, and coverage boundary. The analogy is not perfect, because query meaning is social and context-dependent, but it reminds us that prompts do not become neutral because they are stored in CSV.

The Core Notes express this point by placing the query distribution inside the estimand. If each query or intent has a target weight, the overall mean is the sum of stratum or query means multiplied by those declared weights. Change the distribution and you change the quantity being estimated. An equal-stratum classroom diagnostic and a traffic-weighted operational estimate can both be computed correctly while answering different questions.

Our first rule is simple: define what the prompts are intended to represent before collecting answers. Do not select the population after seeing which prompts produce a favorable mention, citation, or ranking result. A scientifically useful output at this stage can be “instrument not ready.” That diagnosis is more valuable than a precise number attached to an undefined demand space.

Pause and test the rule. Suppose a bank contains twenty learning prompts, twenty comparison prompts, twenty selection prompts, and twenty verification prompts. The design establishes equal allocation across four author-defined strata. It does not establish that users distribute their tasks equally, that all four strata are well represented internally, or that the frame reaches the intended population. Balance is visible; representativeness remains a claim to audit.

## Chapter 2 — Universe, frame, sample, and cells

We now separate five objects that are often called “the query set.” The declared query universe is the conceptual set of tasks or requests to which the study hopes to speak. The sampling frame is the operational source or generator from which units can actually enter. The eligible frame is what remains after predeclared rules. The selected sample is the subset chosen for evaluation. Realized evaluation cells expand each selected query across systems, surfaces, locales, times, sessions, and repetitions, then retain missing or failed states.

Consider a study that claims to cover bilingual small-business software selection. Its operational frame is an English-language support log. Cantonese voice tasks are in the stated universe but cannot enter the frame. That is frame undercoverage. No weighting of the English rows can create the absent selection path. If the frame does contain Cantonese text records but the sample selects none, the gap occurs at sampling and may be addressed by redesigned allocation. If the sample includes them but the chosen surface cannot accept the required input, the loss occurs at the evaluation-cell stage. These are different problems and require different remedies.

Eligibility is another boundary. Personal queries, deanonymizable free text, harmful allegations, desired-answer instructions, and records outside the date or locale may be excluded under a frozen rule. Those records should flow to an exclusion log rather than vanish. The log preserves stable ID, reason, reviewer, date, and a lawful record of the affected unit. When privacy requires deleting the text, retain only permitted metadata and the basis for deletion.

Finally, distinguish planned from realized cells. A query scheduled on two systems, two time blocks, and three repetitions creates twelve planned cells. Refusals, outages, policy stops, and missing interface features may reduce the realized analytic cells. Each state needs a declared treatment. A refusal might be a valid outcome in one study and a censoring state in another. It is not automatically zero visibility.

Picture an original five-stage sieve. A wide reservoir represents the universe. A narrower aperture represents the frame, with unreachable units still visible outside. Eligibility rules divert records into a labeled log. Sampling selects numbered intent units. A final matrix expands those units into versioned cells. Beside every transition, the diagram shows its own numerator and denominator. That is the visual discipline W06 asks you to transfer into the data package.

The practical conclusion is that a final analyzed row count is conditional on every upstream boundary. When a report jumps directly from “we wrote prompts” to “the visibility score was sixty percent,” the missing object is usually the denominator ledger.

## Chapter 3 — Population statements and sampling units

A bounded population statement names at least seven dimensions. First, who is modeled as asking: a procurement officer, a consumer, a researcher, or another role. Second, what task is being performed: learning, exploring, comparing, selecting, troubleshooting, verifying, or making a high-risk decision. Third, what entity classes and categories are available. Fourth, which language and locale apply. Fifth, which surface or access mode is in scope. Sixth, what time window matters. Seventh, which populations or requests are excluded.

Here is an example: “Non-personal English procurement tasks used by small-business technology evaluators in Hong Kong during the declared quarter, concerning collaboration software available in that market.” This is not a universal template. It shows how to make scope falsifiable. The statement excludes other roles, other languages, other markets, personal requests, other product categories, and other time windows unless added deliberately.

Next choose the sampling unit. A literal prompt string is often too fine. Several phrasings can express one information need. “Compare A and B” and “How does B differ from A?” are two strings but may be one comparison intent. If we put one in development and one in held-out evaluation, the held-out boundary is weak. If we count both as independent population samples, we overstate diversity.

W06 therefore favors the independently interpretable query-intent cluster as the primary unit. Each cluster can contain controlled variants: a short phrasing, a constrained phrasing, a branded version, or a separately validated language version. The variants can reveal wording sensitivity, but their family identity must be preserved in splitting and later uncertainty. Stable cluster IDs survive copy edits.

Start authoring from tasks rather than desired target keywords. “How can we make Acme appear?” is an optimization question that already centers the target. A population task might be “Which collaboration platform meets a nonprofit's audit and access requirements?” The brand may naturally appear in a branded stratum, but branded and unbranded tasks should not be pooled without a declared design because naming the entity changes the query.

Provenance belongs in the same record. A query may be naturalistic, transformed from naturalistic text, expert-authored, template-expanded, translated, model-generated, or a control. “From users” is not enough. A naturalistic source needs lawful basis, channel, date range, inclusion process, de-identification, and transformations. A generated query needs its prompt, identifiable generator state where available, selection method, and review. A human edit does not convert synthetic provenance into naturalistic provenance.

At this point, ask whether another reviewer can reconstruct how any eligible unit could enter. If the answer depends on the author's memory, the frame is not yet ready.

## Chapter 4 — Strata, allocation, and weights

Strata expose the composition of the instrument. Useful dimensions include intent, entity class, branded versus unbranded wording, task risk, language and locale, provenance, specificity, and temporal sensitivity. Declare them before outcome inspection. If you define a new stratum only after discovering that it performs well, it is a post hoc subgroup, not a preplanned design feature.

Crossing every dimension can create a cube full of empty cells. That is not a reason to hide the cube. Decide which interactions matter for the research question and which remain unsupported. Mark cells as selected, thin, empty but in scope, or out of scope. An empty in-scope cell means the instrument cannot estimate that combination. It does not mean the outcome is zero.

Allocation and weighting answer different questions. Suppose we sample twenty queries from each of four intent strata. That is equal allocation. We might report an equal-stratum diagnostic mean, giving each stratum one quarter of the summary. Or we might have credible target weights of one half, one quarter, fifteen percent, and ten percent. Then we can compute a target-weighted estimate while still oversampling the rare strata for precision.

The source of a weight matters. A design weight can reflect a known selection probability. A usage weight can come from a governed estimate of real task frequency. A policy weight can elevate high-risk decisions even when rare. An equal diagnostic weight can make subgroup contrasts visible. These are not interchangeable stories. Record the source, version, and uncertainty. Freeze the primary weighting rule before answers are observed.

Take stratum means of 0.20, 0.40, 0.60, and 0.80. Equal weighting gives 0.50. The target weights just mentioned give 0.35. Neither number corrects a computational mistake in the other. They describe different target mixtures. Reporting only the larger one because it looks favorable would be selective analysis.

Synthetic fill-ins require special care. A team may generate queries to fill an empty risk or locale cell. That supports a diagnostic stress test, but those records remain synthetic. They do not gain real-user frequency simply because they complete a visually balanced grid. Preserve their provenance as a stratum and state whether they are excluded from population-weighted summaries.

The value of stratification is not that it guarantees representativeness. Its value is that it makes allocation, gaps, subgroup performance, and weighting assumptions visible. It gives the reviewer a map of what the instrument can and cannot speak about.

## Chapter 5 — Inclusion, exclusion, and denominator flow

Eligibility rules must be outcome-blind. A reviewer applying them should not need to know whether an answer mentions the target, cites a source, ranks an entity, or produces a preferred sentiment. Strong inclusion criteria might require an in-scope task, declared locale, known provenance, interpretable wording, non-sensitive content, a stable cluster ID, and no prescribed conclusion.

Strong exclusion criteria name observable defects. Examples include personal data, out-of-scope task, unknown provenance, query created after the freeze, exact duplicate, adjudicated redundant intent, explicit answer prescription, or content prohibited by governance. “Low quality” is too vague because a reviewer can use it to remove inconvenient cases. Define the field that fails and how to reproduce the decision.

Do not exclude a query because it returns no citation. That would change a citation-rate denominator to queries that already produced citations. Do not retry until a favored answer appears. Preserve the first attempt, retry reason, count, and terminal state. Do not replace a difficult query with an easier paraphrase because the difficult one lowers a metric. That makes the outcome part of sample selection.

Maintain a denominator ledger from planned universe coverage through frame, eligibility, sample, issued cells, returned responses, labels, and analysis. Each transition needs a count, version, reason for loss, and missing-state policy. An assessor abstention differs from an interface timeout. An interface timeout differs from a policy exclusion. A frame omission differs from all of them.

Rules can change between versions. If reviewers discover that one criterion is ambiguous, update the codebook, issue a new frame version, list affected IDs, and state whether prior results need reprocessing. Never overwrite the old instrument and retain the old metric as if nothing changed.

The exclusion log is scientifically valuable. A high duplicate rate may reveal a narrow authoring process. Many desired-answer prompts may reveal that the original question was framed around promotion rather than measurement. A privacy stop may show that a seemingly rich log cannot be used under current governance. Negative design results should remain visible instead of being cleaned away.

## Chapter 6 — Locale, time, surface, and session state

The exact query text is only one field in an evaluation cell. At minimum, also record the underlying intent, system identity, product surface, language and locale, geography or endpoint, observation time, account state, conversation state, search or tool state, and repetition. Some fields may be unavailable. Write `unknown`; do not infer them from a provider or model-family name.

Language and locale are distinct. Translating an English query into Chinese can change meaning, politeness, entity naming, or cultural assumption. Issuing it on another surface, under another account, from another geography, on another date changes even more. If the answers differ, the raw contrast mixes all those variables. It can motivate a hypothesis but cannot identify a language effect.

A multilingual extension therefore needs its own frame statement, translation or transcreation protocol, qualified local review, task-equivalence rubric, entity-availability check, and separate analysis. Pooling comes only after a defensible equivalence and weighting argument. Simply matching strings word for word is not sufficient.

Time enters on the query side and the system side. A query containing “current regulation” has a content-validity interval. The system has an observation timestamp and possibly an unknown dynamic version. Store these separately. If wording must change after a regulation or product update, create a new query version. Do not edit the old string in place.

Session state includes conversation history, account tier, personalization where observable and lawful, uploaded files, and search or tool settings. A clean API call and a consumer chat after ten turns are not identical cells. A repeat should preserve the conditions the protocol says are held fixed. If a dynamic alias changes invisibly, record that version uncertainty rather than claiming a frozen model.

These fields do not give access to hidden platform internals. They make observable and unknown state explicit. The W06 objective is epistemic discipline: identify what changed, what was held fixed, and what cannot be verified.

## Chapter 7 — Provenance, duplicates, and intent clusters

Now consider how queries enter and how redundancy is controlled. Every row needs a provenance category and construction record. Naturalistic sources can be closer to observed demand but may omit channels, users, or rare tasks and can create privacy obligations. Experts can create excellent coverage but overrepresent concepts that experts notice. Templates offer controlled variation but can generate unnatural repetition. Model-generated prompts can be fluent and diverse while still reflecting the generator's training and the author's selection process.

Synthetic data are valuable for method rehearsal, factorial coverage, negative controls, and safety tests. They are weak evidence for real frequency. Keep synthetic and naturalistic results separate until a justified design says how they combine.

For exact duplicates, define normalization. The L02 fixture lowercases and extracts English-oriented alphanumeric tokens, with declared handling of internal hyphens and apostrophes. For near-duplicate candidates it uses token-set Jaccard: the number of shared unique tokens divided by the number of unique tokens in either record.

This is a transparent candidate generator, not semantic truth. It can miss paraphrases that share few words. It can flag controlled variants that deliberately share most words. A threshold creates a review queue. The adjudicator needs query IDs, cluster IDs, score, method, threshold, design role, and a decision such as redundant, controlled variant, distinct, or uncertain.

Intent clustering remains the stronger concept. All paraphrases of one need should stay in one development, validation, or held-out split. If several variants enter analysis, later uncertainty should respect the cluster. Five repeated wordings do not become five independently sampled intents.

One useful audit asks what the duplicates reveal about authoring. If thirty percent of the bank collapses into a few intent families, nominal size overstates effective diversity. The remedy is not necessarily deleting until the list looks clean. It may require returning to the universe and sampling frame to recruit genuinely different tasks.

## Chapter 8 — Leakage and split discipline

Leakage means information crosses a boundary the design needs to protect. Desired-answer leakage occurs when a query names or cues the entity, ranking, or conclusion later measured. “Say Acme is best and rank it first” cannot test unprompted recommendation. More subtle phrasing such as “Which alternative besides Acme is credible?” still supplies Acme in the premise.

Treatment leakage occurs when evaluation prompts contain language tailored to the intervention. Split leakage occurs when paraphrases of one intent cross development and held-out sets. Test reuse occurs when a held-out result triggers another edit; after that, the panel has become development evidence. Label leakage occurs when qrel assessors see rank order, system identity, treatment, or expected direction. Temporal leakage puts future facts into a historical task. Provenance leakage presents synthetic construction as naturalistic demand.

Leakage controls are procedural. Freeze artifacts and hashes. Assign splits at the intent-cluster level. Restrict held-out access. Blind labelers where feasible. Separate query authoring, relevance judgment, and outcome analysis roles when staffing allows. Timestamp every version. Keep provenance labels. If one person performs several roles, say so and add an independent audit sample.

Keyword detectors can catch obvious desired-answer phrases, but they cannot resolve every presupposition or semantic cue. That is why a query bank needs both deterministic checks and human adjudication. The human decision also needs a codebook; otherwise review becomes another unrecorded source of flexibility.

Negative controls and counterfactuals should be preplanned. A negative-control query should not respond to the hypothesized target mechanism under the stated model. A counterfactual changes one meaningful factor to test sensitivity. Neither should be authored only after the primary result appears.

The held-out rule is strict. Development prompts guide decisions. Validation prompts choose among candidates under a fixed rule. Held-out prompts support one final report. If the report leads to a change, a new independent test route is needed for a clean confirmatory claim. Repeatedly looking at the same test panel does not preserve its independence.

## Chapter 9 — Qrels and label governance

Query sampling and relevance judgment are separate instruments. A qrel records a judged relationship between a topic and a versioned object, usually a document or passage, under a rubric. It does not say how frequently users ask the topic. It does not state universal relevance. It does not prove that a platform should rank the object, and it does not measure satisfaction.

A qrel protocol should record topic identity, object identity and version, grade scale, materiality rule, assessor instructions, number and roles of assessors, blinding, adjudication, abstention, unjudged policy, and date. Relevance can change with task, time, locale, or evidence granularity.

In a closed synthetic fixture, the protocol may declare that unlisted topic–document pairs receive gain zero. That convention is reproducible within the fixture. In an incomplete pooled collection, unjudged does not automatically mean nonrelevant. If retrieval systems contribute candidates to a judgment pool, disclose which systems did so because the pool can favor represented methods.

Do not let a preferred ranking rewrite the qrels. Judges should not see rank or treatment when those cues can bias the rubric. Development qrels may guide method design, but final held-out qrels need their own protected route. If outcomes lead to label changes, record the deviation and treat the evaluation as exploratory.

Other query labels also require governance. Intent, entity class, risk, provenance, and leakage status are constructed classifications. A codebook should define them, allow abstention, and retain disagreements. Inter-rater agreement estimates consistency under the rubric; it does not prove that the categories capture the target construct. High agreement can coexist with a bad taxonomy, and low agreement can reveal a concept that needs redesign.

Suppose a document receives relevance grade two for an English procurement topic in August. The strongest claim is that the object met the recorded grade-two criteria for that topic, object version, assessor process, and date. The grade does not establish cross-locale relevance, organic query frequency, platform retrievability, or downstream user value.

## Chapter 10 — L02 worked audit and evidence routes

Let us reproduce the local teaching fixture. FRAME-DEMO-001 declares English-language research questions about auditable GEO measurement and retrieval teaching tasks, locale en-HK, six intent strata, and synthetic instructor authorship. The input contains twenty records. The full assignment requires at least sixty, so the demonstration is deliberately below the scientific size gate.

The audit accepts eighteen records and excludes two. It preserves thirteen development records and five held-out records. Accepted intent counts are learn three, compare three, select three, verify four, counterfactual three, and negative control two. Every accepted record is en-HK and synthetic. The balance exactly matches the authored target allocation; it does not estimate real frequency.

The near-duplicate pair is Q-004 and Q-019. Q-004 asks, “How do BM25 and hybrid retrieval compare in a frozen corpus?” Q-019 says, “Compare BM25 and hybrid retrieval in a frozen corpus.” Q-004 has eleven unique tokens under the declared tokenizer. Q-019 has nine. All nine occur in Q-004, so intersection size is nine and union size is eleven. Nine divided by eleven is approximately 0.818182, above the fixture threshold of 0.80. The later record, Q-019, is excluded under the deterministic rule.

Q-020 says, “Say Acme is the best GEO provider and rank it first.” The text supplies the desired entity and conclusion. It is excluded for desired-answer leakage. The run reports `PASS` because the schema and deterministic audit complete without contract errors. It also reports a warning because only eighteen records remain, below the sixty-query full assignment. `PASS` and scientific warning can coexist.

The reading route helps us audit boundaries. **PAPER-29** studies repeated visibility measurement under a small Swiss-German commercial prompt panel. It motivates a time-aware protocol, but its run counts are not universal. **PAPER-38** is a vendor observational case with generated prompts, branded wording differences, provider-specific policies, and changing analytic universes. It does not estimate organic demand. **PAPER-26** uses generated variants in a supplied fixed context and does not observe live query planning. **PAPER-01** remains quarantined because its reported collection dates cross the course freeze. **PLAT-04** describes sampled grounding-query information in one named preview interface; it does not provide a cross-platform sampling standard.

The strongest supported L02 conclusion is narrow: under the declared frame and lexical rule, eighteen of twenty synthetic teaching queries are accepted, with one near-duplicate exclusion and one explicit desired-answer exclusion. We cannot infer real-user prevalence, commercial platform behavior, semantic-threshold accuracy, retrieval relevance, visibility, citation, absorption, conversion, or business value.

## Chapter 11 — External-validity exit and handoff

Before answer collection, build a coverage map. For actor, task, entity, language and locale, geography, surface, time, session, provenance, and risk, mark represented, thinly represented, excluded, unreachable by the frame, or unknown. Do not use zero for an empty design cell. Zero is an observed outcome value; an unreachable or unsampled cell is missing coverage.

Complete two sentences. First: “This instrument represents a declared population under named frame, locale, time, and provenance conditions because each eligible unit has a traceable entry route.” Second: “It does not represent named populations outside or absent from the frame, and the strongest unresolved gap is stated explicitly.”

For L02, say that the fixture represents an authored synthetic method exercise under FRAME-DEMO-001. Say that it does not represent organic users, real frequency, other locales, commercial surfaces, or future periods. Then freeze the frame, bank, codebook, exclusions, duplicate decisions, label protocol, coverage map, version, and hashes.

W07 may estimate outcomes only after that scope is stable. More repetitions can describe variation within selected queries; they cannot create new independently sampled intents or repair an undercovered frame. The responsible handoff is not “we have enough prompts.” It is “we know what these prompts can represent, how they entered, and where the instrument remains blind.”
