1. The opening distinction: order is not context
Suppose a passage appears first in a ranked list. What has been established? In an inspectable experiment, we may know that a scoring rule placed that passage above the other members of a named candidate pool. We do not yet know that the passage was admitted to a finite model context. Even admission does not establish that the generator relied on its content, that an interface displayed its source, or that a user acted on the answer. W04 begins by protecting these distinctions.
Use five separate event records:
- Candidate entry: the representation belongs to the pool supplied to the reranker.
- Reranked order: the representation receives a position inside that fixed pool.
- Context admission: a passage or excerpt is selected under a declared budget and packing rule.
- Answer use: a controlled analysis finds evidence that content affected the generated response.
- Citation presentation: an interface visibly attributes a source object.
A sixth record, production outcome, covers events such as referral, inspection, sign-up, or purchase. None of these later records should be silently filled from an earlier one. A context allocator might reject a long rank-one passage in favor of several shorter passages. A generator might ignore packed text. A citation resolver might attach a link after generation. An outcome might move because of a concurrent campaign. Each transition needs its own observation or an explicit unknown state.
This vocabulary also disciplines evidence reading. PAPER-17 raises conditional contribution and fair-attribution questions; PAPER-22 examines a relationship between attention and exposure in a bounded research setting; PAPER-18 examines response under evidence conditions. These readings are useful because they make candidate composition, model setting, and outcome definition discussable. They do not jointly reveal one universal production pipeline. Their local catalog status is preprint / venue not confirmed, so the course treats each result as setting-bounded. PLAT-01 is a product-specific official documentation route, useful for the named platform boundary but not for inferring other systems’ hidden ranking or context logic.
Checkpoint A. A response displays a citation to Source S. Which W04 event is directly observed? Citation presentation. Which earlier events may be plausible but remain unobserved without another trace? Candidate entry, reranked order, context admission, and answer use.
2. The experimental contract: change one stage
An interpretable reranking experiment preserves the candidate pool. Let
C_q = {c_1, c_2, ..., c_m}
be the identified candidates for query q. A baseline ordering r_0 and a treatment ordering r_1 may differ, but both must be permutations of the same C_q. Before calculating a metric, verify set equality, candidate count, stable identifiers, and content hashes. If the treatment adds an item, removes an item, changes segmentation, or fetches a newer version, the comparison is no longer pure reranking.
This invariant produces an important ceiling: pure reranking cannot repair a relevant item missing from C_q. Recall at the full candidate depth m is fixed because membership is fixed. Recall at a shallower display cutoff may change when relevant candidates move across that cutoff. NDCG can change because it is rank sensitive. These statements concern the declared qrels and pool, not an external corpus.
A context-packing experiment needs a different invariant. Freeze the ordered passage slate, the passage contents and lengths, the relevance and coverage annotations, and usually the token budget. Then change one allocation rule: for example, add a redundancy penalty, enforce an origin cap, or substitute claim coverage for raw relevance. If both the order and the budget change, the cause of the allocation difference is confounded. A budget ablation may intentionally change B, but the rule and slate must then remain fixed.
The minimal experiment card has three rows:
| State | Reranking comparison | Packing comparison |
|---|---|---|
| Manipulated | ordering function or one reranker feature | one allocation rule, or budget in a separate budget ablation |
| Held fixed | query, candidate IDs, representations, qrels, cutoff | ordered slate, passages, annotations, and all but the named factor |
| Not evaluated | acquisition, generation, citation, outcomes | answer use, citation, outcomes, and hidden commercial behavior |
This is not bureaucratic decoration. It prevents a familiar error: attributing a gain to a reranker when the treatment secretly retrieved a better pool. It also prevents context “optimization” from being evaluated only by how much text fits, without preserving which claims, qualifications, and sources the text covers.
Checkpoint B. Candidate D7 appears only in the treatment run. Can the delta be called a reranking effect? No. Candidate construction, filtering, or representation also changed, even if the treatment component is named a reranker.
3. Candidate identity and comparison units
Candidate identity is more difficult than comparing URL strings. The reranker may operate on documents, passages, chunks, image–text pairs, product records, or synthesized cards. Two rows with the same URL may contain different excerpts. Two passages with different IDs may be exact duplicates. A page updated between runs may preserve its address but change the evidence. A defensible candidate record should include:
- stable candidate and source identifiers;
- parent document identity and version or capture time;
- passage offsets or segmentation rule;
- content hash;
- origin or dependency cluster;
- candidate-generation route and original rank;
- any truncation applied before scoring.
Set equality should therefore be checked over the reranking unit, not merely the source domain. If a baseline scores 500-token passages and a treatment scores 250-token passages, representation changed. If a learned reranker truncates every input at a declared limit, record the actual scored prefix. A candidate can contain the relevant sentence outside that prefix; its poor rerank position is then partly a representation/truncation issue.
Deduplication belongs at a named boundary. Exact content hashes can catch identical records, but syndicated text may differ only in headers or minor edits. Near-duplicate judgments need a documented method and an uncertainty state. Collapsing ten mirrors can reduce source concentration without increasing independent evidence. Conversely, an aggressive deduplicator can erase distinct qualifications that share most wording. Report both the dependency cluster and the individual source identities when the distinction matters.
The L03 fixture gives a clean teaching route. Its hybrid top five forms the pool; the transparent coverage function changes scores only for members of that pool. The validator can reproduce candidate-set identity. The frozen dense column has model identity none; it must not be relabeled as output from A04, Pyserini, RankLLM, or a commercial model. G03 and G04 are candidate implementation routes for separately pinned extensions. Their existence does not change the evidentiary identity of the authored fixture.
Checkpoint C. Two runs contain the same five document IDs, but one scores different passage excerpts from those documents. Candidate identity at the reranker input is not preserved. The comparison mixes representation with ordering.
4. Scores are not meanings
A reranker may be pointwise, pairwise, or listwise. A pointwise function scores each candidate largely in isolation. A pairwise function learns or computes preferences between candidates. A listwise function can depend on the entire slate and its order. The type matters because removing a competitor may change other scores in a listwise system even when their text is unchanged.
Whatever the architecture, record the checkpoint or rule, inputs, candidate depth, truncation, prompt or feature specification, output parsing, tie policy, ordering direction, and deterministic settings where available. Raw scores should not automatically be compared across queries. A score of 0.81 for one query need not mean the same as 0.81 for another. A score is not a calibrated probability of relevance unless calibration has been established for the declared setting. It is never, by arithmetic alone, a probability of truth, source quality, citation, or conversion.
Labels determine what the reranker is being rewarded for. Topical relevance can conflict with evidentiary fitness. A short promotional passage may match a query while lacking a date or condition. A long standards passage may be highly fit for a compliance claim but receive a lower generic relevance score. W04 therefore requires a label card: gain scale, annotation unit, assessor instructions, adjudication, unjudged policy, and known limitations.
Ties also need a rule. Stable candidate ID, original rank, or randomized tie breaking can produce different ordered lists. A deterministic ID tie break helps reproduction but has no semantic meaning. Report it. If scores are rounded for display, evaluate with unrounded values or declare that display rounding defines the tie.
Attention visualizations belong under the same restraint. In an open model, an attention quantity may be inspected under a specified layer, head, tokenization, and aggregation. The result can be a mechanism clue. It does not automatically identify passage utility, causal answer use, displayed citation, or another model’s behavior. PAPER-22 is assigned to sharpen this question, not to license the equation “attention equals exposure” outside its conditions.
5. Reciprocal rank, DCG, and NDCG
Rank-sensitive metrics summarize a judged ordering. They do not replace topic-level inspection.
Reciprocal rank. If the first relevant item occurs at rank r, reciprocal rank is 1/r. It uses a binary relevance decision and ignores later relevant items. Mean reciprocal rank averages the value across the declared topic set. The topic population and missing-result rule must be named.
Discounted cumulative gain. For graded gain g_i at position i, one common convention is
DCG@k = sum_{i=1..k} (2^{g_i} - 1) / log_2(i + 1).
This convention emphasizes high gains and discounts later ranks. Other gain and discount conventions exist. The equation belongs in the metric card.
Normalized DCG. Sort the judged gains into the ideal ordering, calculate IDCG@k, and divide: NDCG@k = DCG@k / IDCG@k. If there is no positive gain, declare the convention rather than silently divide by zero. NDCG is bounded by zero and one under ordinary nonnegative gains, but a score near one says only that the evaluated ordering resembles the ideal ordering defined by those judgments.
Consider gains [3, 0, 2] at ranks one through three. With exponential gain, DCG is 7/1 + 0/log2(3) + 3/2 = 8.5. The ideal gains are [3, 2, 0], so IDCG is 7 + 3/log2(3), approximately 8.8928. NDCG is approximately 0.9558. A different unjudged policy or gain scale changes the result.
At least six fields accompany any reported metric: candidate or corpus scope, relevance unit, gain scale, cutoff, aggregation, and judgment policy. Add tie handling and uncertainty for a stronger record. Macro averaging gives each topic equal weight; a weighted average estimates a different target. Do not select an aggregation after seeing which one favors the treatment.
Metric deltas need direction and stage. “NDCG@3 increased by 0.04 for three authored topics under fixed candidates” is bounded. “The source became more visible” is underspecified because it does not identify the pool, cutoff, event, or downstream boundary.
6. Why macro gains can hide stage failures
A macro mean can rise while some topics regress. The correct response is not to discard the mean but to inspect its composition. For every topic, compare candidate membership, the full before/after order, gains at the evaluation cutoff, and the reason codes for movement. Useful reranking error classes include:
- candidate miss: the relevant unit never entered the pool;
- ordering error: a relevant pool member falls below less relevant members;
- label mismatch: the relevance schema does not capture the task utility;
- representation loss: the scored passage omits the useful evidence;
- truncation loss: relevant evidence exists after the scored prefix;
- duplicate crowding: near-identical candidates occupy scarce top positions;
- unjudged ambiguity: a moved item lacks a reliable judgment;
- tie instability: an arbitrary tie rule produces a material cutoff change.
Each class leads to a different next test. A candidate miss returns to W03 candidate construction. An ordering error invites reranker analysis. Representation loss calls for segmentation work. Duplicate crowding may call for a cluster-aware selector. Unjudged ambiguity calls for assessment, not a confident effect claim.
Negative and unchanged results remain evidence. The L03 fixture’s selected macro NDCG result is unchanged between hybrid and reranked under its default configuration, even though the scoring rule is applied. That is not a failed lesson. It demonstrates why one cannot assume that an added component creates a metric gain. A valid report asks whether any topic order changed, whether the top-k gains changed, and whether the candidate pool remained fixed.
PAPER-17, PAPER-18, and PAPER-22 can motivate additional mechanism or allocation questions, but a package analysis must keep its own estimand. A result about conditional contribution cannot be silently converted into a result about retrieval. A result under fixed evidence composition cannot establish organic source acquisition. A result relating an internal quantity to exposure in one setting cannot be asserted for a closed service without separate evidence.
7. Context packing as constrained evidence allocation
After reranking, a system may have more candidate text than the context budget permits. Packing selects passages or excerpts, their order, and sometimes their truncation. A transparent ledger gives each passage:
- a stable ID and source-origin cluster;
- token cost
l_iunder a declared tokenizer or supplied authored count; - relevance value
v_i; - coverage indicators
e_icfor material claimc; - redundancy relation
u_ijto other passages; - required qualification or contradiction flags;
- selection state, selected span, order, and exclusion reason.
One teaching formulation chooses binary z_i to reward relevance and claim coverage while penalizing redundant pairs, subject to sum(z_i l_i) <= B. This is a design model, not a description of a platform. Coefficients encode preferences. Increasing a diversity term can remove the strongest single source; decreasing redundancy can omit a necessary corroboration. The impact must be inspected rather than presumed desirable.
Four measurements should remain separate:
- Relevance total: sum or distribution of the declared passage values.
- Material-claim coverage: proportion of preregistered claim requirements covered by at least one selected passage.
- Origin diversity: count or distribution of dependency clusters, with a declared independence interpretation.
- Redundancy: exact or judged overlap under a specified rule.
Source count alone is not diversity. Five domains can repeat one syndicated statement. Redundancy is not always waste: two independent sources may corroborate a material fact. Claim coverage is not truth: an annotation can be wrong. A context ledger should preserve these limitations.
Greedy packing by rank is easy to reproduce but can perform poorly under variable lengths. A long top passage may consume the budget while covering only one claim. Knapsack-style optimization can improve an authored objective but introduces coefficient and annotation sensitivity. Whichever rule is used, preserve rejected candidates and reason codes. A polished context should not erase its allocation history.
8. A valid packing ablation
An ablation identifies what one component contributes under a frozen setting. Consider a baseline that admits passages in reranked order until the next passage exceeds B. The treatment uses the same ordered slate, budget, lengths, and claim labels but skips a passage if its evidence content is near-duplicate of an already selected passage. Only the redundancy rule changes.
The report should include:
- exact baseline and treatment selected IDs in order;
- total authored token cost and unused budget;
- material claims covered and missing;
- origin clusters represented;
- redundant pairs admitted or avoided;
- exclusions with reasons;
- sensitivity to uncertain redundancy labels;
- all downstream events marked not evaluated.
Do not change B while calling the comparison a redundancy ablation. Do not let the treatment reranker see different passages. Do not update labels after observing the selection. If passage lengths are estimated rather than produced by the intended tokenizer, label them authored costs.
A separate budget curve can hold the rule fixed and evaluate several preregistered budgets. Its conclusion is conditional: under this ledger and rule, coverage changed at these budgets. It does not reveal the budget of an opaque service. The apparent “knee” of a curve is sensitive to passage segmentation and the chosen claim requirements.
Checkpoint D. Baseline uses B=180; treatment uses B=240 and a redundancy penalty. Is the delta attributable to redundancy control? No. Budget and rule changed together. Split the work into two comparisons.
9. Answer use, citation, and production are separate studies
A packed context is an input record. To study answer use, generate under controlled conditions and compare outputs with claim-level annotations. A leave-one-passage-out change can show conditional influence under that context, model, prompt, and decoding configuration. It cannot prove unique authorship or global importance. Shapley-style contribution likewise depends on the player definition, coalition distribution, and utility. PAPER-17 is valuable precisely because those assumptions deserve inspection.
Citation requires another schema. A displayed link might be selected jointly with answer text, resolved from metadata afterward, or attached by an interface layer. Citation correctness asks whether a cited passage supports the adjacent claim. Completeness asks whether material claims receive support. Neither can be derived from reranker rank alone.
Production outcomes require an observation or experiment at the user–task level. Context inclusion is not a visit. A citation is not a conversion. If referral is measured, define the interface, population, exposure, time window, missing events, concurrent changes, and privacy controls. W04 does not conduct that study.
For a closed answer surface, an analyst may record the query, visible response, citation objects, time, locale, account state, and collection method. Unless the system exposes more, candidate membership, ranking features, context order, and answer-use cause remain unknown. An open sandbox teaches mechanisms by exposing variables; it does not authorize reverse-engineering claims about a different system.
PLAT-01 may document named product behavior and explicit boundaries. Its first-party authority is claim-relative and platform-specific. It cannot turn an observed citation into a known reranker trace or establish a cross-engine allocation rule.
10. Writing the bounded W04 conclusion
A defensible conclusion contains six parts:
- Setting: fixture, corpus snapshot, query set, and unit.
- Comparison: manipulated component and held-fixed invariants.
- Measurement: gain scale, cutoff, aggregation, judgment policy, and packing budget if relevant.
- Result: topic-level evidence before the macro summary.
- Failure or limitation: at least one candidate, label, representation, or packing limitation.
- Ceiling: downstream and production events not evaluated.
Example:
In the ten-document, three-topic L03 synthetic fixture, the transparent coverage reranker operated only over each fixed hybrid top-five pool. Candidate identity was preserved. Under the declared qrels and cutoff, selected macro NDCG@3 remained unchanged at the reproduced precision, although score values changed. This result concerns ordering in the authored fixture; context packing, answer generation, citation, and production behavior were not evaluated by the L03 script.
For a packing case:
In the five-passage authored ledger, the redundancy-aware rule used the same passage order, 180-token budget, costs, and coverage labels as greedy rank packing. It replaced one near-duplicate with a complementary passage and increased material-claim coverage from two of four to three of four under the supplied annotations. The case does not establish answer use, citation, or behavior on a named platform.
Notice the absence of an inflated adjective such as “better” without an object. The reranker may improve NDCG and worsen latency. The packer may improve claim coverage and reduce corroboration. State the measured axis.
11. Discussion prompts, summary, and exit test
Discussion prompts
- When is a document ID insufficient to prove fixed candidate identity?
- Can a pure reranker change Recall at candidate depth? Can it change Recall at a shallower cutoff? Explain both.
- Which NDCG assumptions are hidden when a chart displays only one decimal score?
- When should redundancy be retained rather than penalized?
- What would a leave-one-out generation test add after context admission, and what would it still fail to establish?
- How should PAPER-22 be discussed without converting a bounded attention/exposure result into a universal platform explanation?
- What product-specific claim can PLAT-01 support, and which cross-engine claim remains outside its ceiling?
Summary
Reranking is ordering conditional on a candidate pool. Preserve candidate identity before assigning a delta to the reranker. Rank-sensitive metrics interpret judged order under explicit conventions; they do not measure answer use. Context packing is a constrained allocation whose budget, passage unit, coverage, redundancy, and provenance choices must remain visible. A fixed-order, fixed-budget ablation can identify a change in the authored allocator. It cannot identify a closed platform’s rule. Packed, used, cited, and outcome events remain distinct.
Exit test
Write two sentences. Sentence one must name a fixed candidate pool, reranking change, metric, cutoff, and topic-level result. Sentence two must name the strongest downstream unknown. Then annotate each noun with its unit and each claim with its evidence source. If the result cannot be reconstructed without oral explanation, revise it before submission.