# W04 Slide Script — Ranking and Context Selection

## Production conventions

- **Format:** 16:9 academic lecture deck using deep purple, warm ivory, graphite, and restrained gold. Color never carries a state by itself.
- **Typography:** minimum 30 pt body, 44 pt title, and 24 pt table labels. Equations receive a plain-language reading.
- **Visual authorship:** every visual below is an original instructional construction. Do not reuse a paper figure, product screenshot, repository interface, or platform logo.
- **Boundary grammar:** solid outline means observed in the fixture; dashed outline means inferred; hatched fill means not observed; a double border marks a held-fixed object.
- **Accessible production:** preserve the exact alt text, include data tables in notes, announce rank changes verbally, and never make animation necessary for comprehension.

## Slide 01 — From candidate to context

**On-screen text**

> W04 · Ranking and context selection  
> Essential question: Why does an eligible source enter the context?

Footer: “Candidate ≠ rank ≠ packed ≠ used ≠ cited.”

**Visual specification**

Five labeled gates run left to right: Candidate, Reranked, Packed, Used, Cited. The first three have solid outlines; the last two are hatched. A small production-outcome box sits beyond a gap. No gate resembles a commercial interface.

**Speaker notes**

Open by asking what rank one proves. In an inspectable run it proves an ordering event inside a named pool. It does not prove context admission or answer influence. Read the footer slowly and define each verb. Explain that W04 studies fixed-pool ordering and finite-budget allocation. Generation and citation are downstream studies, and a closed product may expose none of the intermediate gates.

**Teaching check**

Learners write the one event directly observed when a public answer shows a citation.

**Alt text**

“Five separate gates connect candidate entry to visible citation, with downstream gates hatched to show that they remain unobserved in the opening example.”

## Slide 02 — Outcomes leave audit artifacts

**On-screen text**

By the end, you can:

1. Prove fixed candidate identity.
2. Calculate a rank-sensitive metric.
3. Isolate one context-packing factor.
4. Diagnose the first loss stage.
5. Bound open and closed claims.

Artifacts: equality proof · metric card · allocation ledger · stage record · boundary memo.

**Visual specification**

Five numbered rows pair each outcome with a document-shaped artifact. Thin vertical lines connect the rows to a final “reconstruct without oral explanation” test. Labels and numbers establish association without icon dependence.

**Speaker notes**

Frame the week as an evidence-production exercise. A verbal assertion that candidates were fixed is insufficient; the equality proof needs IDs and versions. A reported NDCG needs its gain, cutoff, tie, and judgment policy. A packing result needs every rejected passage and reason. The final criterion is whether another analyst can reconstruct the comparison from the artifacts alone.

**Teaching check**

Ask which artifact would reveal that a treatment quietly introduced a sixth candidate.

**Alt text**

“Five learning outcomes align with five auditable artifacts and converge on a reconstruction test.”

## Slide 03 — The six-event vocabulary

**On-screen text**

| Event | Observable record |
|---|---|
| Candidate entry | pool membership |
| Reranked order | position and score |
| Context admission | span, order, budget |
| Answer use | controlled influence evidence |
| Citation presentation | visible source object |
| Production outcome | user–task event |

**Visual specification**

A vertical event ledger uses six distinct row numbers. Beside each is a miniature record with fields, not an illustrative product view. Between rows, conditional arrows include the word “may.” The last three records use hatched backgrounds.

**Speaker notes**

Contrast observability with logical possibility. A cited source probably had some path into an attribution process, but the visible citation alone does not expose the candidate list, ranking features, packed span, or causal influence. “Used” requires an answer-level design such as a controlled ablation and claim analysis. Outcomes require a user-level record. Never backfill hidden records from the final screen.

**Teaching check**

Give the sentence “Source S was used because it was cited.” Learners identify the unsupported transition.

**Alt text**

“Six event rows list the distinct record needed for candidate entry, reranking, packing, answer use, citation, and a production outcome.”

## Slide 04 — Stage isolation is the experiment

**On-screen text**

**Reranking:** freeze candidates; change order.  
**Packing:** freeze order and budget; change one allocation rule.  
**Budget ablation:** freeze order and rule; change only budget.

Three columns: Manipulated · Held fixed · Not evaluated.

**Visual specification**

Three horizontal experiment strips. Each has one gold movable block, several double-bordered fixed blocks, and a hatched downstream area. A legend spells out movable, held fixed, and not evaluated using words and line styles.

**Speaker notes**

Explain why names are not evidence. A component called “reranker” can also fetch, filter, or segment. The comparison earns the reranking label only when candidate identity is preserved. A redundancy ablation earns its label only when rank, budget, costs, and annotations remain fixed. If budget and rule change together, separate them into two experiments.

**Teaching check**

Treatment changes the ranking model and candidate depth. Is the delta a pure reranking effect? Learners answer no and name the second changed factor.

**Alt text**

“Three experiment strips show exactly one manipulated factor, multiple fixed factors, and downstream states excluded from evaluation.”

## Slide 05 — Candidate identity is more than a URL

**On-screen text**

Candidate record:

- stable candidate ID
- source and capture version
- passage offsets or segmentation
- content hash
- origin cluster
- truncation before scoring
- original recovery route and rank

**Visual specification**

One URL card branches into two differently bounded passage cards with distinct hashes and offsets. A second pair of different URLs converges on one dependency cluster because their passages are near duplicates. Both patterns have explanatory labels.

**Speaker notes**

A treatment can preserve URLs while changing the reranker inputs. Different excerpts, updated content, or truncation create different candidates. The reverse problem also occurs: different domains may mirror one syndicated passage, so counting URLs can overstate evidentiary diversity. Set equality must operate over the actual scored unit and its version, not a convenient address field.

**Teaching check**

Two runs contain the same document IDs but use different chunk sizes. Which stage besides ordering changed?

**Alt text**

“The same URL can produce different passage identities, while different URLs can belong to one duplicated-origin cluster.”

## Slide 06 — The fixed-pool proof

**On-screen text**

Required checks:

`IDs_0 = IDs_1`  
`count_0 = count_1`  
`hash(id)_0 = hash(id)_1`  
`gain(id)_0 = gain(id)_1`

Then—and only then—interpret order.

**Visual specification**

Two five-row lists sit side by side. Candidate cards appear in different orders, but dotted identity lines connect matching IDs and hashes. A green textual badge reads “same members, same versions”; a separate arrow marks changed positions.

**Speaker notes**

Set equality alone catches additions and removals but not duplicated IDs, content drift, or changed judgments. Verify uniqueness, counts, versions, and gains. The gain equality is an evaluation invariant, not necessarily part of the deployed reranker. If relevance labels are revised after results are seen, create a new analysis version and preserve the original.

**Teaching check**

Ask why two sets can be equal while one list contains a duplicated row. What additional check is needed? Candidate count and uniqueness.

**Alt text**

“Two differently ordered five-item lists have matching IDs, hashes, and gains connected across columns, proving candidate preservation before rank interpretation.”

## Slide 07 — What reranking cannot recover

**On-screen text**

Relevant universe in fixture: `{A, B, C, D}`  
Fixed candidate pool: `{A, B, X}`

No permutation can introduce C or D.

At full pool depth, membership recall is invariant.

**Visual specification**

A large relevant-set oval contains A through D. A double-bordered pool rectangle encloses A, B, and X. Three small permutations of the same pool appear below. C and D remain outside every permutation.

**Speaker notes**

Pure reranking can move A and B but cannot recover C and D. At the full three-item pool depth, the relevant membership count is fixed. At a shallower cutoff, order can change which relevant members appear. This is why Recall@3 can move inside a fixed top-five pool while Recall@5 remains constant.

**Teaching check**

Learners explain, in one sentence each, how Recall at cutoff three may change while Recall at candidate depth five does not.

**Alt text**

“Repeated permutations of one candidate pool never include the relevant items that were absent from the pool.”

## Slide 08 — Scores require an identity card

**On-screen text**

Record:

- pointwise, pairwise, or listwise
- model/checkpoint or transparent rule
- input and truncation
- prompt/features and output parsing
- candidate depth
- tie policy and score direction
- deterministic settings

Score ≠ calibrated probability.

**Visual specification**

A score column is covered by a translucent identity card listing seven fields. Without the card, the column ends in a question mark. A separate scale shows 0.8 on two queries with a broken equality sign between them.

**Speaker notes**

A decimal does not explain its production. Pointwise and listwise functions have different dependencies. Scores may not be comparable across queries and may shift when the pool changes. The L03 reranker is an authored coverage function; the frozen dense run has model identity `none`. Do not rename either using A04, G03, G04, or a product name.

**Teaching check**

Ask whether score 0.81 for Query A must indicate more relevance than 0.74 for Query B. What evidence would calibration require?

**Alt text**

“A reranker score becomes interpretable only when paired with a complete model or rule identity card; identical decimals across queries are not equated.”

## Slide 09 — Labels define the task

**On-screen text**

Gain card:

`0` not useful  
`1` partly useful  
`2` directly useful  
`3` central audit contract

Also declare: unit · assessor rule · adjudication · unjudged policy · limitations.

**Visual specification**

Four gain rows use numbers, text, and progressively thicker borders. Alongside them, one passage receives two possible labels under different task cards: topical match versus evidentiary fitness.

**Speaker notes**

Relevance is not a natural property detached from a question. A promotional paragraph may match a topic and still be poor evidence for a dated compliance claim. A long technical passage may be evidentially strong but weak under a generic topical label. NDCG rewards the declared labels, so the label card belongs beside every metric result.

**Teaching check**

Learners name one passage that could be topically relevant but unfit for the claim being answered.

**Alt text**

“A four-level gain scale and its annotation contract show that the same passage can receive different judgments for different declared tasks.”

## Slide 10 — Reciprocal rank sees the first hit

**On-screen text**

First relevant rank `r` → `RR = 1/r`

Examples:

- rank 1 → 1.00
- rank 2 → 0.50
- rank 5 → 0.20
- no relevant result → declared zero convention

Blind spot: later relevant evidence.

**Visual specification**

Four short ranked strips highlight the first relevant item with a numbered bracket. Later relevant items are outlined but faded, with a label that reciprocal rank ignores them.

**Speaker notes**

Reciprocal rank answers when the first item crossing a binary relevance rule appears. It does not reward a second or third useful source, graded relevance, diversity, or material-claim coverage. Mean reciprocal rank also needs a declared topic set and missing-result convention. Use it only when the first useful hit matches the estimand.

**Teaching check**

Compare lists `[relevant, not, not]` and `[relevant, relevant, relevant]`. What does reciprocal rank report, and what does it miss?

**Alt text**

“Four ranked strips map first-relevant positions to reciprocal-rank values and explicitly fade later relevant items that the metric ignores.”

## Slide 11 — DCG rewards gain near the top

**On-screen text**

`DCG@k = Σ (2^{g_i} − 1) / log2(i + 1)`

For gains `[3, 0, 2]`:

`7/1 + 0/log2(3) + 3/2 = 8.5`

Convention alert: gain and discount must be stated.

**Visual specification**

Three position columns display gain, transformed gain, discount, and contribution. Bar heights show contributions, while printed numbers provide the exact data. A small side card names cutoff three.

**Speaker notes**

Read the equation in words: sum each transformed gain divided by a logarithmic position discount through cutoff k. Explain that exponential gain makes the difference between grades two and three larger than between zero and one. Other conventions exist. A displayed score without its convention cannot be exactly reproduced.

**Teaching check**

Learners calculate the contribution of gain two at rank three and explain why it equals 1.5 under this convention.

**Alt text**

“A three-position table decomposes DCG into transformed gain, logarithmic discount, and numeric contribution, totaling 8.5.”

## Slide 12 — NDCG normalizes to the judged ideal

**On-screen text**

Observed gains: `[3, 0, 2]` → `DCG = 8.5000`  
Ideal gains: `[3, 2, 0]` → `IDCG = 8.8928`  
`NDCG@3 = 0.9558`

Near one means near the *declared judged ideal*.

**Visual specification**

Two aligned three-slot lists compare observed and ideal gains. Curved arrows move gain two from rank three to two. A fraction bar displays 8.5000 over 8.8928. A boundary caption rejects answer-quality interpretation.

**Speaker notes**

Normalization supports comparison across topics with different gain totals, under the same convention. It does not make the labels objective or convert the result into answer faithfulness. Specify what happens when no positive gain exists. Keep unjudged documents distinct from judged zero unless the fixture explicitly adopts the simplification.

**Teaching check**

Ask whether NDCG 0.96 proves that the generated answer is correct. Learners name the missing generation and claim-evaluation stages.

**Alt text**

“Observed and ideal gain orders produce an NDCG fraction of 8.5 divided by 8.8928, while a caption limits interpretation to judged ranking.”

## Slide 13 — The W04 five-candidate trace

**On-screen text**

Baseline: `C1, C2, C3, C4, C5`  
Reranked: `C2, C4, C1, C5, C3`  
Gains: `C1=1, C2=3, C3=0, C4=2, C5=1`

Same IDs · same hashes · same gains.

**Visual specification**

Two ordered stacks show the five IDs. Lines trace each candidate’s movement. C2 and C4 move upward, C3 downward, and C1/C5 shift modestly. Each card prints its gain and matching hash prefix.

**Speaker notes**

Introduce the authored worked case. The baseline top-three gains are one, three, zero. The reranked top-three gains are three, two, one. Before celebrating the order, point to the equality proof. This case is intentionally clean and synthetic; the scores have no external population meaning alone.

**Teaching check**

Learners identify which positively judged candidate remains below the reranked cutoff and why reranking cannot say whether it belonged in a larger pool.

**Alt text**

“Five fixed candidates move between baseline and reranked stacks while retaining the same gains and version hashes.”

## Slide 14 — Calculate before interpreting

**On-screen text**

Baseline:

- `NDCG@3 = 0.576666646`
- `Recall@3 = 0.50`
- `Recall@5 = 1.00`

Reranked:

- `NDCG@3 = 1.000000000`
- `Recall@3 = 0.75`
- `Recall@5 = 1.00`

**Visual specification**

A two-column metric table is followed by three delta rows. The unchanged full-depth recall row uses an equals sign, not a flat green line. A note identifies the four positive-gain members.

**Speaker notes**

Explain the shallow recall increase as movement across rank three. Full-depth recall is invariant because membership is fixed. The large NDCG delta was constructed for transparent arithmetic, not as an expected empirical effect. Report the fixture, topic, labels, cutoff, and candidate depth whenever quoting a number.

**Teaching check**

Ask a learner to give the numerator and denominator for reranked Recall@3 and for Recall@5.

**Alt text**

“A metric table compares baseline and reranked NDCG and recall, highlighting that Recall at full candidate depth remains exactly one.”

## Slide 15 — Macro means require topic stories

**On-screen text**

Inspect before averaging:

- candidate miss
- ordering error
- label mismatch
- representation or truncation loss
- duplicate crowding
- unjudged ambiguity
- tie instability

Keep gains, regressions, and unchanged topics.

**Visual specification**

Seven diagnostic drawers sit beneath a macro score card. Three topic cards—gain, unchanged, regression—are placed into different drawers with reason labels. No drawer is colored as inherently good or bad.

**Speaker notes**

A mean can increase while a material topic regresses. Diagnosis begins with the first boundary supported by evidence. A missing candidate returns to retrieval. A present but low-ranked candidate is an ordering issue. An omitted qualification can be representation or label mismatch. Do not use the macro score to erase an adverse case.

**Teaching check**

A relevant passage is in the pool but below cutoff after treatment. Which error classes are plausible, and what record distinguishes them?

**Alt text**

“A macro score opens into seven diagnostic categories and three preserved topic outcomes: gain, unchanged, and regression.”

## Slide 16 — Context is a finite allocation

**On-screen text**

Passage ledger fields:

`ID · cost · relevance · claim coverage · origin · redundancy · selected span · exclusion reason`

Constraint: `Σ selected_cost ≤ B`

Packed order is a new object.

**Visual specification**

An ordered candidate rail feeds passage blocks of different lengths toward a 180-unit context bar. Each block carries claim letters and origin codes. Rejected blocks remain visible below with reason tags.

**Speaker notes**

Packing can select entire passages, excerpts, and order. A long high-ranked item can crowd out several shorter complementary items. The ledger must retain rejected candidates; otherwise, the polished context hides the selection process. Costs require a tokenizer identity or, as in this case, an explicit authored-unit label.

**Teaching check**

Ask why “three sources fit” is not reproducible without passage units, costs, and a budget.

**Alt text**

“Differently sized passage blocks compete for a fixed 180-unit context bar, with rejected blocks and exclusion reasons preserved.”

## Slide 17 — Four packing measurements, not one score

**On-screen text**

Report separately:

1. relevance value
2. material-claim coverage
3. dependency-origin diversity
4. redundancy or corroboration

Source count ≠ independent evidence.  
Coverage label ≠ truth.

**Visual specification**

A selected two-passage context fans out into four small meters with printed values. Two apparent domains merge into one dependency cluster in the diversity meter. A redundancy pair is labeled “may be corroboration.”

**Speaker notes**

Collapsing all objectives into an undocumented quality score hides tradeoffs. Multiple domains can repeat one press release. Redundancy can waste budget, but independent corroboration can be valuable. Claim coverage depends on an annotation schema and is not itself a truth test. Report each axis and its uncertainty.

**Teaching check**

Learners explain why removing every repeated claim could reduce evidence quality even when it increases unique-claim count.

**Alt text**

“One packed context is evaluated on four separate axes, while two domains merge into one dependency cluster and redundancy remains potentially useful corroboration.”

## Slide 18 — Greedy packing in the worked case

**On-screen text**

Fixed order: `C2(95), C4(75), C1(80), C5(55), C3(40)`  
Budget: `180`

Greedy selection: `C2 + C4 = 170`  
Coverage: `{A,B}` = `2/4`  
Origins: one dependency cluster  
Redundant pair: `(C2,C4)`

**Visual specification**

The 180-unit bar contains a 95-unit C2 block and a 75-unit C4 block, leaving a labeled 10-unit gap. Claim letters A and B appear on C2; A repeats on C4. C1, C5, and C3 remain below.

**Speaker notes**

The baseline follows rank and selects any passage that fits. C2 and C4 consume 170 units. Although they have different candidate IDs, the origin ledger places them in one mirrored dependency cluster, and C4 adds no new material claim under the frozen annotations. Do not call this failure yet; it is the baseline allocation under a declared rule.

**Teaching check**

Ask which evidence fields, besides length, establish that C4 is redundant for this task.

**Alt text**

“C2 and C4 fill 170 of 180 units, repeat claim A, cover only A and B, and come from one dependency cluster.”

## Slide 19 — One-rule packing ablation

**On-screen text**

Only change: skip a same-cluster near duplicate with no new material claim.

Selection: `C2 + C1 = 175`  
Coverage: `{A,B,C}` = `3/4`  
Origins: two clusters  
Redundant pairs: zero

Not changed: order · budget · costs · labels.

**Visual specification**

The fixed candidate rail remains above. C4 moves to a rejected row with reason code. C2 and C1 fill the same 180-unit bar. A side checklist uses double borders around every held-fixed element.

**Speaker notes**

The treatment skips C4 and admits C1. The authored coverage increases, but D remains missing. The result is conditional on the redundancy label. If C4 contains a unique qualification, the treatment may be worse. We have isolated a rule effect in this fixture, not found a preferred platform strategy.

**Teaching check**

Learners name the exact manipulated factor and four held-fixed factors.

**Alt text**

“A redundancy rule replaces C4 with C1 under the same order and budget, increasing labeled coverage from two to three of four claims.”

## Slide 20 — Sensitivity is part of the result

**On-screen text**

Revisit if:

- C4 has a unique qualification
- origin clusters are uncertain
- costs use a different tokenizer
- passage segmentation changes
- requirement D is weighted as mandatory

Preserve the adverse interpretation.

**Visual specification**

Five assumption switches connect to the treatment result. Flipping any switch leads to a separate versioned branch, not an overwritten score. One branch displays “unresolved” rather than a forced winner.

**Speaker notes**

An ablation is only as credible as its frozen labels. Sensitivity analysis asks whether plausible annotation or representation changes reverse the decision. Never edit the original ledger after seeing the result. Version a revised case. A mandatory limitation claim could make both allocations inadequate despite the numerical coverage change.

**Teaching check**

Ask which change is a label sensitivity and which is a representation change: making D mandatory versus shortening C1.

**Alt text**

“Five explicit assumptions branch into versioned sensitivity analyses, including an unresolved outcome rather than a forced positive conclusion.”

## Slide 21 — Packed does not mean used

**On-screen text**

Packed passage → controlled generation study → claim-level change

Possible evidence:

- leave-one-passage-out comparison
- fixed prompt/model/decoding
- claim segmentation and support labels
- repeated runs if stochastic

Still not: authorship · global importance · automatic citation.

**Visual specification**

A packed-context box enters two parallel generation conditions, full and leave-one-out. Their answer claims are compared in a small diff ledger. A boundary wall blocks arrows to authorship and universal platform behavior.

**Speaker notes**

Answer use needs a separate controlled design. Removing one passage and observing a claim change can indicate conditional influence in that context. It can be confounded by equivalent information in other passages and stochastic generation. PAPER-17 helps ask contribution questions, but player definitions and utility functions remain part of the estimand.

**Teaching check**

Learners state one conclusion a leave-one-out change supports and one it does not.

**Alt text**

“Full-context and leave-one-out generation conditions feed a claim-difference ledger, while boundaries block authorship and universal-mechanism claims.”

## Slide 22 — Used does not mean cited

**On-screen text**

Citation evaluation asks:

- Is a source object displayed?
- Does the cited passage entail the adjacent claim?
- Are material claims covered?
- Is the source fit for the claim?

Rank alone answers none of these.

**Visual specification**

An answer with three atomic claim cards connects to two source-passage cards using solid support, dashed partial support, and one unresolved edge. A separate presentation layer adds visible citation badges after a boundary line.

**Speaker notes**

An interface can attach citation objects through a stage separate from generation. A high-ranked passage may not be cited; a cited passage may not support the adjacent claim. Citation correctness and completeness require claim segmentation and exact passage evidence. W05 develops this work; W04 keeps the boundary visible.

**Teaching check**

Ask why a visible link cannot reveal its source’s candidate rank or whether the generator relied on it.

**Alt text**

“Atomic answer claims connect to evidence passages with different support states, while visible citation badges occur in a separate presentation layer.”

## Slide 23 — Reading ceilings and closed systems

**On-screen text**

- **PAPER-17:** conditional contribution and attribution questions
- **PAPER-22:** bounded attention/exposure mechanism route
- **PAPER-18:** evidence-composition extension
- **PLAT-01:** named-product documentation only

Local status: papers are preprints / venue not confirmed.  
Closed candidate pool, reranker, and context order: `unknown` without evidence.

**Visual specification**

Four source cards sit beneath a horizontal “claim ceiling” line. Above the line are prohibited universal-mechanism boxes. To the right, a closed-system silhouette contains three hatched unknown fields and one visible citation record outside it.

**Speaker notes**

Reading is not license to merge estimands. Each paper supports claims bounded by its system, candidates, model, and metric. PLAT-01 is authoritative for the named documentation, not for cross-engine mechanisms. On a closed surface, record what is visible and leave hidden intermediate variables unknown. An analytical diagram remains a hypothesis unless instrumented.

**Teaching check**

Learners rewrite “the platform chose this source because it had the highest attention” as a bounded observation plus a mechanism hypothesis.

**Alt text**

“Four bounded source routes remain below their claim ceilings, while a closed-system silhouette preserves candidate, reranker, and context fields as unknown.”

## Slide 24 — The bounded W04 conclusion

**On-screen text**

A complete conclusion names:

1. setting and unit
2. fixed pool or fixed order
3. manipulated factor
4. metric, cutoff, and policy
5. topic-level result and limitation
6. downstream events not evaluated

Exit: “Where did the evidence first disappear?”

**Visual specification**

Six sentence blocks form a narrowing template. Each block has a labeled input field and an evidence icon. At the bottom, a stage trace offers candidate miss, ordering loss, packing exclusion, use unknown, and citation unknown as distinct outcomes.

**Speaker notes**

Close by reading a bounded result: in the authored five-candidate pool, IDs and versions remained fixed; the declared reranker changed top-three metrics. In the separate fixed-order packing case, one redundancy rule changed allocation under 180 units. Neither result evaluates answer use, citation, or production. Ask learners to locate the first evidence-supported loss stage rather than invent a hidden cause.

**Teaching check**

Learners submit two sentences: one reconstructable result and one strongest downstream unknown, each with its unit and evidence source.

**Alt text**

“A six-part conclusion template ends in a stage-specific trace that distinguishes candidate, ordering, packing, use, and citation states.”
