Explain

W06 · 5 h 15 min

Queries as a measurement instrument

Essential question

What distribution do the prompts represent?

What you should be able to do

  1. Define a query universe.
  2. Construct a stratified sample.
  3. Prevent desired-answer leakage and near duplication.
PrerequisitesBasic sampling vocabulary.Ten candidate queries without target-brand wording.
Builds on

Basic sampling vocabulary. · Ten candidate queries without target-brand wording.

Investigates

What distribution do the prompts represent?

Feeds

L02 query-set engineering.

Arrive with a prepared artifact

Reading route
Core PAPER-29/PAPER-38; Extend PAPER-01/PAPER-26.
Viewing route
Open the complete text-first lecture packageSix-minute query-leakage and near-duplicate audit. The current equivalent is notes, slide script, worked case, and no-video transcript; no recording is claimed.
Readiness check
Assign ten seed queries to intent and entity strata; flag sensitive or desired-answer wording.
Bring
Draft population statement and exclusion rules.

Concepts, assumptions, and boundary

A query list implies a population

Visibility estimates are conditional on the prompts selected. A convenient list scraped from one tool or invented around a desired outcome rarely represents a defensible population. A query universe names who asks, what intent is expressed, which entities and locales are in scope, how prompts arise, and what time window matters. Strata make the design inspectable; they do not guarantee representativeness unless the sampling frame and weights are justified.

Prompt construction can leak the answer

Queries that include desired entities, treatment language, or outcome wording can mechanically inflate mention or citation rates. Near duplicates over-weight a narrow phrasing family and make nominal sample size misleading. Counterfactual prompts, negative controls, multilingual extensions, and excluded-query logs help diagnose these risks. Synthetic queries must be labeled and analyzed separately from naturalistic data because fluency does not establish ecological validity.

Inspect the mechanism or evidence structure

Conceptual modelQuery-universe sampling cube
Query-universe sampling cubeThe cube shows that a flat query list samples combinations of intent, entity, locale, and time. Empty cells and overrepresented cells are labeled; a table provides identical strata and counts.IntentEntityLocaleTimesampled cells + weights
Figure design. A three-axis cube for intent, entity class, and locale/time, with selected cells, unrepresented cells, and weights shown directly. Empty and overrepresented cells define the inference boundary.
Long description

The cube shows that a flat query list samples combinations of intent, entity, locale, and time. Empty cells and overrepresented cells are labeled; a table provides identical strata and counts.

One action, one feedback state

Action

Allocate a fixed sample budget across query strata.

Feedback

The tool reports coverage, expected variance, and which population claims become unsupported.

Accessible alternative

Four predefined allocation plans and calculations are provided as a table.

Open the repeated-measurement explorer

Produce a reviewable intermediate file

Task
Build and audit a versioned query bank with at least four intent strata.
Inputs
Sampling-frame template, taxonomy, deduplication script, synthetic seed set.
Timebox
75 minutes
Intermediate file
queries-v1.csv, codebook, duplicate report
Stop condition
Exclude personal or sensitive queries unless an approved de-identification route exists.

Why each controlled source is here

Complete, retrieve, and revise

Checkpoint
L02 query-set engineering.
Reflection
State which users or intents the query bank cannot represent.
Revision
Resolve duplicates and leakage before any response collection begins.
Low-compute route
Use the supplied seed set and deterministic similarity report; no embedding model is required.

Five retrieval questions

01Why does query deduplication affect inference?

Near duplicates overweight a phrasing family, reduce effective sample diversity, and can narrow the population represented.

02What is desired-answer leakage?

Prompt wording that directly names or cues the entity, treatment, or conclusion being measured.

03How should multilingual extensions be handled?

Declare language and locale strata, validate translation/cultural equivalence, and analyze them separately before pooling.

04What is the evidence boundary for W06?

The papers provide measurement and field-study cases; their prompt distributions are not automatically suitable for another domain.

05What must you submit or revise after this week?

L02 query-set engineering. Resolve duplicates and leakage before any response collection begins.

Continue in the practice package