# W12 Lecture Notes — Optimization under Competing Queries

**Essential question:** Which gain survives a portfolio of intents?  
**Evidence status:** Conceptual method plus a deterministic, authored teaching matrix. No live platform, model, user, ranking, citation, or business outcome was observed.

## 1. Replace “optimize the page” with a declared decision problem

Optimization is not a synonym for improving everything. It is a selection procedure over a permitted set of revisions. A complete statement names the decision unit, candidate set, target query portfolio, outcomes, objective directions, constraints, budget, uncertainty method, and choice rule. Without those objects, “best content” is not a research result. It is an unbounded preference.

Begin with a versioned source state (x_0). A candidate (x_j) is a typed revision: target component, operation, affected claim IDs, evidence IDs, owner, cost, authorization, and rollback. The permitted set contains only candidates whose factual message and required qualifications remain supported. A new unsupported proposition is not a stronger version of an old candidate; it belongs outside the permitted set until evidence and approval exist.

W12 asks a portfolio question because one source state can serve different intents. A learner may want a definition, a buyer may compare alternatives, a reviewer may verify a method, and a skeptical reader may ask for contrary evidence. These tasks can prefer different placement, density, and detail. A short summary may improve quick orientation while removing context needed for verification. A dense limitations block may help skeptical use while imposing cost elsewhere. Conflict is therefore expected, not an error to be averaged away.

Write the decision before seeing outcomes:

> Among the seven frozen candidates, select at most one revision for the fictional local evidence page, using five declared intent strata and an eight-point edit budget, after non-compensable integrity gates, with a frozen weighted rule, a lower-tail condition, uncertainty threshold, tie rule, and rollback path.

This wording does not promise release. “No candidate” and “retain control” are valid actions. It also prevents the evaluated candidate set from silently expanding after a preferred answer appears.

## 2. Declare the query portfolio and weights before outcomes

A query portfolio approximates a target demand distribution; it is not automatically the population of all users. The population statement should identify the organization, page, language, locale, date window, task boundary, and exclusions. Queries are grouped by underlying intent before splitting or weighting, so paraphrases of the same fact cannot be scattered across development and evaluation cells as if independent.

The W12 teaching portfolio has five fictional strata:

| Stratum | Weight | Intended information need |
|---|---:|---|
| definition | 0.30 | Establish meaning, entity, and scope. |
| compare | 0.25 | Contrast alternatives on declared dimensions. |
| select | 0.20 | Support a bounded choice without persuasion shortcuts. |
| verify | 0.15 | Recover method, date, evidence, and limitations. |
| counterfactual | 0.10 | Find failure cases, exceptions, or reasons not to choose. |

The weights sum to one. They are governance inputs, not discovered constants. They might represent a declared service mandate, a frozen survey estimate, or a policy allocation. Their provenance must be recorded. If an analyst changes the counterfactual weight after observing its loss, the analyst has changed the target policy. That can be a legitimate sensitivity analysis, but it cannot replace the primary preregistered result.

Several checks precede scoring. Does each stratum have a definition? Are excluded desired-answer prompts documented? Are weights fixed before candidate outcomes? Are locales and dates stable? Are safety-critical or minority intents protected by an explicit rule rather than a tiny weight? Could a single source page plausibly serve all strata, or should the design route people to separate evidence objects? These questions connect W12 to W06: optimization cannot repair an undeclared or contaminated query frame.

## 3. Separate the objective vector from hard constraints

For feasible candidate (j), define a direction-aware vector:

\[
\mathbf{J}_j = [g_{j,d}, g_{j,c}, g_{j,s}, g_{j,v}, g_{j,f}, -k_j],
\]

where the five (g) terms are changes from control in the declared stratum outcome and (k_j) is edit cost. Higher is better in every displayed coordinate because cost is negated. A vector preserves the trade-off pattern. It does not require pretending that a gain of 0.05 in one stratum is naturally exchangeable with a 0.05 loss in another.

Now define hard constraints separately. W12 requires pass states for factual integrity, legal permission, authorization, safety, privacy, and essential accessibility, plus cost no greater than eight edit points. These values are Boolean gates or direction-specific thresholds, not reward components. The correct sequence is:

1. reject a candidate outside the permitted change space;
2. reject any hard-gate or budget failure;
3. compare objective vectors only among survivors;
4. apply the predeclared portfolio and choice rules; and
5. require accountable approval before any release.

Never write a score such as `0.7 visibility + 0.2 factuality + 0.1 accessibility` if low factuality or inaccessible essential evidence should block release. Such a sum allows a large proxy gain to purchase permission to be wrong or unusable. A hard constraint is intentionally non-compensable. This is not merely conservative arithmetic; it expresses who has authority to expose others to risk.

Budget also needs a unit. The eight points in the teaching case are authored effort units, not currency or time. A real study might track editor hours, token change, component count, API calls, money, or review capacity. Different budgets should not be added without a conversion rule. The control costs zero and remains an eligible action, which is why it can remain on a Pareto frontier even when several revisions have positive gains.

## 4. Freeze the outcome matrix and preserve every adverse cell

The complete authored W12 matrix is below. Each outcome is a change from the local control on a fictional normalized score. `half-width` is a frozen symmetric teaching interval half-width, not an estimate from observations. All candidates except C3 meet the eight-point cost limit; C4 fails factual-integrity and safety gates. Other integrity gates pass in this compact fixture.

| ID | Candidate | definition | compare | select | verify | counterfactual | cost | half-width | Gate status |
|---|---|---:|---:|---:|---:|---:|---:|---:|---|
| C0 | control | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0 | 0.000 | pass |
| C1 | evidence proximity | 0.05 | 0.04 | 0.02 | 0.06 | -0.01 | 3 | 0.018 | pass |
| C2 | answer-first limitations module | 0.08 | 0.01 | -0.02 | 0.07 | 0.00 | 5 | 0.020 | pass |
| C3 | C1+C2 bundle | 0.11 | 0.05 | -0.01 | 0.10 | -0.02 | 9 | 0.025 | **budget fail** |
| C4 | persuasive append | 0.03 | 0.10 | 0.12 | -0.08 | -0.05 | 4 | 0.030 | **factual and safety fail** |
| C5 | decorative headings | 0.02 | 0.01 | 0.00 | 0.00 | -0.01 | 4 | 0.015 | pass |
| C6 | definition anchors | 0.02 | 0.02 | 0.01 | 0.04 | 0.01 | 2 | 0.017 | pass |

The negative values are load-bearing. C1 has the highest admissible weighted mean but harms the counterfactual stratum. C2 helps definition and verification but reaches the allowed loss floor on selection. C5 has a weighted interval crossing zero. C3 looks numerically attractive but is not affordable. C4 appears attractive on comparison and selection but is inadmissible before any scalar score is considered.

A transparent matrix retains nulls, failures, and rejected candidates. Hiding C4 would erase evidence that a score can be gamed by an impermissible revision. Hiding C5 would make optimization look more certain than the frozen case. Hiding the negative counterfactual cell would turn a portfolio trade-off into a false universal gain.

## 5. Compute weighted means, lower tails, and uncertainty

The primary scalar summary is the frozen weighted mean

\[
\bar g_j = 0.30g_{j,d}+0.25g_{j,c}+0.20g_{j,s}+0.15g_{j,v}+0.10g_{j,f}.
\]

For C1, the calculation is

\[
0.30(0.05)+0.25(0.04)+0.20(0.02)+0.15(0.06)+0.10(-0.01)=0.0370.
\]

The seven weighted means are C0 `0.0000`, C1 `0.0370`, C2 `0.0330`, C3 `0.0565`, C4 `0.0410`, C5 `0.0075`, and C6 `0.0200`. The highest raw number belongs to C3, yet C3 violates budget. C4 is also higher than C1 but fails factual and safety gates. Neither enters the release comparison.

The symmetric intervals are constructed as weighted mean plus or minus the frozen half-width. C1 is `[0.0190, 0.0550]`; C2 `[0.0130, 0.0530]`; C3 `[0.0315, 0.0815]`; C4 `[0.0110, 0.0710]`; C5 `[-0.0075, 0.0225]`; C6 `[0.0030, 0.0370]`; and C0 `[0.0000, 0.0000]`. These intervals are arithmetic fields in a teaching fixture. They do not have sampling coverage because no samples generated them.

The preregistered portfolio floor requires no stratum change below `-0.02`. The minimum worthwhile lower interval endpoint is `0.015`. After hard filtering, C1 meets both. C2 meets the stratum floor but its lower endpoint is only `0.013`; C6 is robustly positive per stratum but its lower endpoint is `0.003`; C5 is compatible with a null under the authored interval. Null is not failure of the exercise. It is a legitimate disposition: retain control or gather a better-designed evidence set without post-hoc factor addition.

## 6. Determine Pareto dominance only among comparable feasible options

Candidate A dominates candidate B under the W12 convention if both pass hard gates and budget, A is at least as good as B on all five stratum gains, A costs no more, and A is strictly better on at least one coordinate. The convention must be written because other analyses might compare different objectives or treat cost only as a constraint.

C1 dominates C5. Its changes `(.05,.04,.02,.06,-.01)` are at least C5’s `(.02,.01,.00,.00,-.01)` in every stratum, and its cost `3` is lower than `4`. C6 also dominates C5. Therefore C5 cannot be rationally selected under this vector, even before its interval crosses zero.

Among the hard-feasible set `{C0,C1,C2,C5,C6}`, the nondominated frontier is `{C0,C1,C2,C6}`. Control remains because no revision matches its zero cost. C1 and C6 trade higher broad gains against cost and counterfactual protection. C2 trades definition and verification gains against selection loss and cost. A point on the frontier is not automatically safe, approved, or optimal. Frontier membership only states that no compared survivor is unambiguously better under the declared coordinates.

C3 and C4 should remain visible on a plot but outside the feasible region. Crosshatching, an `X`, and explicit text labels distinguish them without relying on color. Do not call them dominated: infeasibility is a different reason for exclusion. First gate, then compare dominance. This keeps “high score but prohibited” from appearing as a near-winner.

## 7. Treat the weighted sum as policy, then test policy sensitivity

A weighted sum resolves trade-offs by declaring exchange rates. It answers, “Which feasible candidate best serves this particular allocation?” It does not answer, “Which candidate is naturally best?” The weights embed stakeholder judgments about whose intents count and how much. A low-weight group can bear a material loss while the mean rises.

The primary W12 policy first filters hard gates and budget, then requires the `-0.02` per-stratum floor, then requires a lower interval endpoint of at least `0.015`, and finally maximizes the frozen weighted mean. It selects C1. If eligible candidates differ by at most `0.002`, the tie rule chooses lower cost, then fewer changed components, then control. If no candidate clears every rule, retain C0 and report an inconclusive portfolio result. If a selected candidate later fails any gate, revert without rescoring.

Now run a preregistered policy-sensitivity analysis. A maximin policy selects the candidate with the largest minimum stratum gain among hard-feasible candidates. The minima are C0 `0.00`, C1 `-0.01`, C2 `-0.02`, C5 `-0.01`, and C6 `0.01`. Maximin selects C6, not C1. Nothing about the outcome matrix changed; the policy changed. The contrast makes governance visible. The weighted rule prioritizes more total declared benefit; maximin prioritizes the worst-served stratum.

The correct report retains both: “C1 is selected by the primary weighted and meaningful-effect rule; C6 is selected by the robust maximin sensitivity rule.” Calling C1 universally superior would erase this disagreement. Decision-makers must state why the primary policy has authority and who bears C1’s counterfactual loss.

## 8. Use single-component ablation without overstating mechanism

An ablation removes one component while holding the rest of a defined construction fixed. C3 is the bundle of C1 evidence proximity and C2’s answer-first limitations module. The matrix supplies both component candidates, enabling a descriptive two-removal calculation even though the bundle itself exceeds budget.

Removing proximity from C3 yields C2. The conditional weighted difference is `0.0565 - 0.0330 = 0.0235`. Removing the answer-first module yields C1, for `0.0565 - 0.0370 = 0.0195`. The bundle’s departure from simple additivity is `0.0565 - 0.0370 - 0.0330 = -0.0135`. At the stratum level, the interaction vector is `(-0.02, 0.00, -0.01, -0.03, -0.01)`.

This is not proof that proximity causally contributes `0.0235` on a real system. The values are authored. Even in an empirical factorial design, conditional contribution depends on the other component, assignment, units, measurement, and interaction assumptions. C3’s budget failure also means the ablation can inform redesign but cannot rescue the bundle for release.

When a clean one-factor comparison is impossible, the design should not pretend. Options include a factorial design with enough support for interactions, a sequential design with a frozen order and stopping rule, or an explicitly bundled estimand. Each changes the question. W09 and L05 establish the discipline of one interpretable factor; W12 adds the possibility that useful portfolio choices require comparing components and conflicts without hiding co-interventions.

## 9. Preregister choice, tie, null, conflict, failure, and rollback rules

A decision rule should be executable by someone who has not seen the desired answer. The W12 card freezes:

- **eligibility:** all factual, legal, authorization, safety, privacy, essential-accessibility, and cost gates pass;
- **portfolio floor:** no stratum below `-0.02`;
- **minimum worthwhile condition:** weighted interval lower endpoint at least `0.015`;
- **primary choice:** greatest frozen weighted mean among survivors;
- **tie:** within `0.002`, choose lower cost, then fewer changed components, then control;
- **null:** an interval crossing zero is reported as compatible with no portfolio gain;
- **conflict:** any negative stratum is named in the release record even when allowed;
- **failure:** no survivor means retain C0 rather than relax rules;
- **rollback:** any later integrity, accessibility, authorization, or protocol failure restores the control hash; and
- **deviation:** a changed weight, threshold, candidate, or matrix version creates a new analysis, never an overwritten primary result.

These rules prevent researcher degrees of freedom. Without them, almost any candidate can be made attractive by changing weights, dropping a stratum, choosing a different cutoff, or reporting only a favorable metric. A complete record preserves the action taken, the losing groups, every failed gate, the uncertainty, and the strongest alternative policy.

## 10. Interpret nulls, conflicts, and constraint violations as results

An optimization study is informative when no candidate survives. It may reveal that the edit budget is unrealistic, the factor is weak, the outcome is noisy, the page cannot serve all intents, or a desired gain depends on impermissible claim drift. Those findings should change the design, not the reporting standard.

The W12 candidates illustrate distinct dispositions. C5 is feasible but dominated and compatible with a null. C2 is feasible and nondominated but misses the minimum worthwhile lower endpoint; it is not “a failed page.” C6 is feasible, nondominated, and favored under maximin, but the primary weighted rule does not select it. C3 has the largest raw weighted mean yet is budget-inadmissible. C4 has attractive cells but fails factual integrity and safety. C1 is selected yet retains a counterfactual loss and a local uncertainty range.

Use precise language. Say “excluded for budget,” not “performed poorly.” Say “factual and safety gate failure,” not “lower quality.” Say “not selected under the primary policy,” not “ineffective.” Say “compatible with a null under the authored interval,” not “proved no effect.” These distinctions keep optimization from becoming a leaderboard divorced from evidence and governance.

## 11. Read PAPER routes as bounded design evidence

PAPER-03 is the core multi-query conflict route. It supports studying heterogeneous revision requirements under a content budget and motivates conflict-aware coordination. It does not establish that its framework, stability metrics, or reported gains transfer to every query distribution, source type, or production engine.

PAPER-31 is the core feature-level multi-objective route. It motivates separating interpretable feature configurations from surface realization and examining more than one objective. Its benchmark results remain conditional on the evaluated engines, candidate process, feature definitions, quality measures, and reporting. It does not validate the W12 numbers.

PAPER-04 extends the comparison to output-order control and provides a useful manipulation-risk discussion. Its conditions do not justify appending steering material to a live source, and the course does not reproduce its system claims. PAPER-27 provides a confidence-decay and deterministic-agent proposal for conceptual critique; it does not reveal hidden confidence state or routing logic in a named commercial system. PAPER-30 motivates proposal archives, critics, diverse strategies, cost accounting, and component ablation; an adaptive proposer still lacks release authority and can optimize a proxy while damaging facts.

The strongest common lesson is methodological: define candidate space, query distribution, comparison, metric, constraints, and transfer boundary. None of the five routes supplies a universal strategy or production effect claim. The attached TeX source and local Notes inform structure only; no external figure or prose is copied into this package.

## 12. Reproduce the frozen decision and write the evidence ceiling

The validator reads `manifest.json`, recalculates every weighted mean and interval, verifies that weights sum to one, applies hard gates and the cost cap, computes dominance, enumerates the frontier, applies the primary and maximin policies, verifies the ablation differences, and reruns the L05 auditor in a temporary directory. It also locks the five narrative hashes and exact seven-file package. Standard-library validation makes the teaching decision inspectable.

Determinism begins with identity, not with a plot. The manifest gives the matrix a version, lists strata in an exact order, records each decimal as a JSON number, names every direction, and freezes the cost unit. The calculation retains full stored precision until display. Candidate IDs are stable, so a label change cannot silently attach outcomes to the wrong revision. Expected results are derived from the matrix during validation rather than copied from a chart. The original heatmap and frontier specifications are presentation layers; neither is an authoritative data source.

A reproduction record should distinguish three failure families. An **identity failure** occurs when a narrative hash, candidate ID, input file, or L05 fixture hash differs. A **contract failure** occurs when weights no longer sum to one, a hard gate is missing, a budget has no direction, or a slide lacks its accessible alternative. A **decision failure** occurs when recalculation changes feasibility, dominance, the frontier, the selected candidate, or ablation result. The validator stops on all three. It does not quietly substitute defaults, discard malformed rows, renormalize weights, or accept extra package files.

Reproduction also needs a human-readable calculation trace. For each candidate, preserve the five products used in the weighted mean, the unrounded total, minimum stratum, interval endpoints, gate failures, and dominance witness. For the primary decision, show every sequential filter rather than only the winner. This makes a crucial distinction inspectable: C3 and C4 are removed for inadmissibility, C5 is removed by dominance and the meaningful-effect rule, and C2 and C6 remain legitimate alternatives that lose under one declared policy. An implementation that sorts raw means before filtering can yield the wrong release while producing numerically correct dot products.

Finally, deterministic reproduction does not create external validity. Re-running an authored matrix perfectly can prove only that the specified teaching calculation was preserved. An empirical study would additionally need registered sampling, assignment or comparison evidence, measurement validation, missingness rules, uncertainty appropriate to its unit structure, versioned systems, and independent review. It would also need to preserve new nulls and harms rather than replacing the frozen result. W12 treats computational repeatability as one layer of quality control, not as a shortcut to scientific or production truth.

The valid conclusion is deliberately narrow:

> In the frozen W12 synthetic matrix, after applying non-compensable integrity gates and the eight-point budget, C1 is selected by the preregistered weighted, floor, and minimum-worthwhile rule. C6 is selected by the maximin sensitivity policy. C1 loses `0.01` in the counterfactual stratum; C5 is dominated and interval-compatible with a null; C3 and C4 are inadmissible. These arithmetic results do not measure a platform response or establish a universal content strategy.

A validator PASS means the package is structurally complete and the local arithmetic is reproducible. It does not mean the authored values are empirical, the source page is accessible, the papers have been independently replicated, learners will understand the lesson, or any production outcome will occur. The responsible final action remains conditional: prepare a human-reviewed, accessible, authorized local candidate, preserve control and rollback, and obtain new evidence before making a stronger claim.
