Prevent desired-answer leakage and near duplication.
PrerequisitesBasic sampling vocabulary.Ten candidate queries without target-brand wording.
Builds on
Basic sampling vocabulary. · Ten candidate queries without target-brand wording.
Investigates
What distribution do the prompts represent?
Feeds
L02 query-set engineering.
02
Before class
Arrive with a prepared artifact
Reading route
Core PAPER-29/PAPER-38; Extend PAPER-01/PAPER-26.
Viewing route
Open the complete text-first lecture packageSix-minute query-leakage and near-duplicate audit. The current equivalent is notes, slide script, worked case, and no-video transcript; no recording is claimed.
Readiness check
Assign ten seed queries to intent and entity strata; flag sensitive or desired-answer wording.
Bring
Draft population statement and exclusion rules.
03
Explanation
Concepts, assumptions, and boundary
A query list implies a population
Visibility estimates are conditional on the prompts selected. A convenient list scraped from one tool or invented around a desired outcome rarely represents a defensible population. A query universe names who asks, what intent is expressed, which entities and locales are in scope, how prompts arise, and what time window matters. Strata make the design inspectable; they do not guarantee representativeness unless the sampling frame and weights are justified.
Prompt construction can leak the answer
Queries that include desired entities, treatment language, or outcome wording can mechanically inflate mention or citation rates. Near duplicates over-weight a narrow phrasing family and make nominal sample size misleading. Counterfactual prompts, negative controls, multilingual extensions, and excluded-query logs help diagnose these risks. Synthetic queries must be labeled and analyzed separately from naturalistic data because fluency does not establish ecological validity.
04
Primary visual
Inspect the mechanism or evidence structure
Conceptual modelQuery-universe sampling cube
Figure design. A three-axis cube for intent, entity class, and locale/time, with selected cells, unrepresented cells, and weights shown directly. Empty and overrepresented cells define the inference boundary.Long description
The cube shows that a flat query list samples combinations of intent, entity, locale, and time. Empty cells and overrepresented cells are labeled; a table provides identical strata and counts.
05
Interactive check
One action, one feedback state
Action
Allocate a fixed sample budget across query strata.
Feedback
The tool reports coverage, expected variance, and which population claims become unsupported.
Accessible alternative
Four predefined allocation plans and calculations are provided as a table.