Explain

W05 · complete online lecture

Generation, Source Use, Citation, and Absorption

Essential question

Was a source merely displayed, or did it shape the answer?

A structurally complete authored draft

Structurally complete authored package draft · human review and timed pilot pending

3,001online lecture words
4,044transcript words
24specified slides
90+90scheduled contact minutes

What the package must enable

  1. Segment generated answers into material atomic claims while preserving scope, quantities, conditions, dates, and qualifications.
  2. Separate source presence, source use, claim absorption, displayed attribution, claim–passage entailment, and answer correctness.
  3. Resolve citation objects to versioned source identities, exact passages, and provenance origins before scoring support.
  4. Compute citation correctness and completeness with declared units, denominators, weights, unresolved states, and eligibility rules.
  5. Preserve independent labels and adjudicate disagreements without deleting pre-adjudication evidence.

Planned 90-minute evidence sequence

W05 seminar plan
MinutesSegmentLearner evidence
0–20 minOpening contrast, generation condition, and effective-context recordFive-event classification and six-field condition card
20–50 minAtomic claims, citation objects, correctness, and completenessSegmented claim, source record, and denominator worksheet
50–75 minAbsorption, controlled source removal, and citation/support matrixAbsorption label, counterfactual design, and failure code
75–84 minBounded paper and platform readingPortable idea and non-portable claim per source
84–90 minExit test and studio handoffTwo-sentence bounded conclusion

75-minute core artifact route inside a 90-minute studio

The remaining 15 minutes are a declared delivery margin for setup, accessible pacing, questions, recovery, and submission packaging; they are not unplanned teaching content.

W05 core studio plan
MinutesActivityStop or redirect condition
0–25 minFreeze identities and independently segment/code five claimsStop on unresolved source identity, claim boundary, or passage locator.
25–50 minExchange labels and preserve six disagreement rowsDo not reveal rationales before the second label or erase pre-adjudication states.
50–70 minAdjudicate selected claims and recalculate support coverageKeep correctness distinct from completeness and unresolved distinct from zero.
70–75 minState the evidence ceilingBlock a mechanism claim that crosses from the open fixture into a closed platform.

Read, teach, inspect, or download

Complete lecture text

This HTML is generated from the controlled Markdown source. Source SHA-256: ed457c77bebcf1e8f3475bb5f074073f5b33688094cc8c3e38b37840b22f5080.

On this page 12 sections

1. The opening distinction: a visible source is not a complete explanation

Suppose a generated answer contains a link to source S. What has been observed? At minimum, a citation object was displayed and resolved to S under recorded conditions. That observation does not, by itself, show that S was retrieved, inserted into the effective context, read by the generator, responsible for any claim, accurately represented, or selected for a defensible reason. Those are separate propositions with different evidence requirements.

W05 uses a six-part separation:

  1. source presence — S occurs in a logged candidate set, context, or visible-source list;
  2. source use — changing S while preserving the rest of a controlled condition changes a declared answer property;
  3. claim absorption — preregistered, source-distinctive information from S appears faithfully in the answer;
  4. displayed attribution — the interface presents a citation resolved to S;
  5. claim–passage entailment — the resolved passage supports the attached atomic claim; and
  6. answer correctness — the claim is defensible against the relevant evidence and scope, including evidence outside S.

No adjacent pair is automatically equivalent. A system may list S without using it. Several redundant passages may make removal of S appear inert even though its information was available. An answer may reproduce a distinctive datum from S without displaying S. A citation resolver may attach S after text generation. A passage may faithfully say something that is factually wrong or obsolete. Conversely, a correct background fact may appear without any demonstrable dependence on S.

This separation changes the research question. Instead of asking “Was S used?”, ask which event is being measured, with what unit, under which system state, and against which alternative explanation. The resulting answers are narrower but more useful because they direct the next diagnostic test.

2. Record the generation condition before interpreting the answer

A generated answer is conditional on more than a user query. In an inspectable system, record the query, conversation state, effective context, system and tool instructions, model and version, decoding or sampling state, and time. In a closed surface, some of these variables remain unavailable. Record the observable interface and mark unobserved components as unknown; do not fill them with assumed architecture.

The effective context is the exact material available to generation after retrieval, reranking, deduplication, truncation, and packing. A search result page, citation list, or publisher dashboard is not necessarily that context. The difference matters. If a source was a retrieved candidate but removed during packing, a generation-stage explanation is premature. If an interface adds citations after sentence production, a mismatch can arise at attribution without implying that generation ignored the passage.

Use a condition card with at least:

FieldExample recordWhy it matters
observation identityOBS-W05-001prevents screenshots from becoming anonymous evidence
surface and configurationnamed surface, explicit version if exposeddefines the measured object
query and prior stateexact string plus conversation historycontrols answer condition
effective contextordered passage IDs and hashes, or unobservedseparates selection from generation
outputraw answer plus structured citationspreserves the event
time and localeISO date-time, language, location/account statebounds changeable behavior

For a closed surface, “effective context: unobserved” is a valid record. It limits inference but does not invalidate the visible observation. A responsible protocol preserves the difference between absence of evidence and evidence of absence.

3. Segment material claims without destroying their meaning

Citation scoring begins only after the answer is segmented into assessable claims. A material atomic claim is the smallest proposition that can receive a support judgment while retaining its qualifiers. Atomic does not mean stripping away conditions. Consider: “Plan K retains records for 30 days except for accounts under the legacy agreement.” At least two predicates are present: a general retention duration and a scoped exception. If the first is assessed while the exception disappears, the annotation has changed the answer’s meaning.

Each claim record should preserve:

  • exact answer span;
  • normalized proposition;
  • subject and predicate;
  • quantity and unit;
  • population or entity scope;
  • time or version scope;
  • condition, comparator, and exception;
  • epistemic or normative modality, such as “may,” “requires,” or “causes”;
  • materiality or weight; and
  • response identity and span offsets.

Segmentation is a measurement decision. Two coders may disagree about whether a conjunction contains one claim, two claims, or a claim plus qualification. That disagreement changes denominators for completeness and unsupported-claim rates. Therefore freeze the segmentation guide before looking at method comparisons, double-code a subset, preserve the original spans, and report how adjudication changed the count.

Three tests are helpful. First, can one clause be true while another is false? If yes, split unless a qualification would be orphaned. Second, could different sources support the clauses? If yes, split and retain their relationship. Third, would a reader be misled if the condition were evaluated separately? If yes, preserve a linked claim–qualification unit. The goal is reproducible judgment, not maximal fragmentation.

4. Resolve citation objects before judging support

A visible citation label is not yet an evidence passage. Resolution records the displayed anchor, target URL, redirects, canonical identity, page version or access time, exact cited passage when available, and provenance origin. It also records failure: unavailable page, blocked access, changed content, ambiguous anchor, or unresolved redirect. Unresolved items remain unresolved; they are not automatically incorrect or correct.

The source object and passage object serve different roles. Source quality is claim-relative and concerns whether the source type is fit for the proposition. Entailment concerns whether a particular passage supports the proposition. A low-authority marketing page can entail the meta-claim “the company page states that its method ranks first,” yet it does not establish that the method really ranks first. A high-authority standard may be irrelevant to a product-effect claim. A scientific paper may support a finding within its experiment while not supporting a claim about a later platform version.

Provenance origin is also essential. Five pages that syndicate one press release are five URLs but perhaps one evidence origin. Counting them as independent corroboration launders duplication into apparent consensus. A citation audit should cluster mirrors, canonical copies, and derivative summaries when the relationship is knowable. Disagreement about origin becomes an uncertainty field.

The claim–citation edge should be created only when the interface places a citation at the claim, or when a published resolver rule maps a citation to it. Global source lists require an explicit placement policy. “A page appears somewhere in the citations” is source presence, not proof that the page supports every claim in the answer.

5. Entailment, completeness, and correctness use different denominators

Let material claims be a_1 ... a_J, citation edges be E, claim weights be w_j, and S_jk equal one when resolved passage k supports claim j under the codebook. A weighted citation-entailment score is:

sum over displayed claim–citation edges of w_j S_jk / sum over displayed edges of w_j.

This asks: among attached citation relationships, how much is supported? A weighted completeness score is:

sum over claims of w_j times an indicator that at least one attached citation supports claim j / sum over material claims of w_j.

This asks: among claims that require evidence, how much receives adequate support? A response can score high on the first and low on the second if its few citations are correct but many material claims are uncited. It can show many citations yet score poorly on both if they attach to the wrong claims.

Citation correctness is sometimes used as a broader name for link validity, placement, and entailment. The package avoids relying on that label alone. A metric card must say which component is tested. Similarly, answer correctness is not citation entailment. To assess answer correctness, evaluate the proposition against an appropriate evidence set, source currency, applicable scope, and contradictions. If S supports a stale value, the answer may have high S-entailment but low current correctness.

Do not silently remove unresolved citations from denominators. Report at least the eligible claim count, displayed-edge count, resolved-edge count, support-positive count, unresolved count, weighting rule, and treatment of qualifications. Where severe claims matter more, define weights before seeing system outputs and show the unweighted result as a sensitivity analysis.

6. The four-cell citation–support matrix

Cross two observable judgments: whether a citation is displayed for a claim and whether an appropriate resolved passage supports the claim.

Passage supports claimPassage does not support claim
Citation displayedsupported citationunsupported or wrong-source citation
No citation displayedsupported but unattributed claimunsupported and unattributed claim

The lower-left cell needs care. Support may be found during an external audit even though the answer did not attribute it. This does not prove the system used that passage. It shows that the proposition can be supported and that attribution is incomplete under the declared policy. The upper-right cell includes misplaced anchors, partial passages attached to compound claims, contradictions, and citations whose pages discuss the topic without entailing the proposition.

Add source quality as a separate axis. A supported citation from a source unfit for the claim should not be converted to “unsupported” if it really entails the wording; instead record entailment and quality separately. Then add answer correctness. This prevents a single label from hiding whether the problem is placement, support, authority, or factual validity.

The matrix is a diagnostic interface, not a total quality score. A safe workflow assigns a disposition to every problematic cell: accept, qualify, replace source, move citation, split claim, correct claim, remove claim, or escalate. An unresolved case remains visible in the release ledger.

7. Absorption requires source-distinctive content and alternatives

Absorption concerns whether information associated with S appears faithfully in the answer. The term should not be assigned merely because S is cited or because answer and page share generic words. Build a preregistered set of claims whose content is distinctive enough to identify, verify each against S, and check plausible alternative origins.

Use a three-level label:

  • none — no verified distinctive content from S is present;
  • plausible — aligned content appears, but common knowledge, duplicated origins, or alternative sources prevent a specific attribution;
  • source-distinctive — a verified datum, formulation, structure, or qualification associated with S appears, and the stated alternative-source check materially narrows competing origins.

Even the strongest label is not automatically causal source use. In a closed system, the model may know the fact parametrically or obtain it from an unseen duplicate. A source-distinctive label is an operational trace under the registry and checks, not proof of memory provenance or authorship.

An absorption rate needs the same discipline as citation metrics. If C_S is the preregistered set and each claim has weight omega_c, report the weighted fraction faithfully present. State who judged faithful presence, how contradictions and omitted qualifications are handled, and how many claims were unresolved. If a model judge assigns labels, calibrate it against humans across languages, claim types, source lengths, and difficulty strata.

Absorption can also concern answer structure, but structural similarity is especially vulnerable to generic templates. A “definition–benefit–steps” sequence rarely identifies one source. Treat such patterns as plausible unless the structure is genuinely distinctive and alternatives are examined.

8. Source use is a controlled contribution claim

Source use is stronger than absorption because it concerns contribution to an outcome. In an open, controlled generator, compare a full context C with a condition C-minus-S that removes S. Hold the query, instructions, model, decoding policy, other passages, context budget, ordering policy, and repetition plan as stable as possible. Then define the outcome before running the comparison: verified claim coverage, qualification preservation, answer correctness, or another bounded utility.

A difference can still be difficult to interpret. Removing S may shift every later passage position. The context assembler may insert a replacement. Token budget may change. Stochastic decoding creates variation. Another source may contain the same evidence. A careful design declares whether the empty space remains, whether a replacement is allowed, and how order is preserved. Use repeated paired runs when randomness is present.

A leave-one-source-out contrast estimates contribution under the constructed context, not global importance. It does not prove authorship, training-data contribution, or a production platform mechanism. A near-zero effect can mean non-use, redundancy, weak measurement, or insufficient repetitions. A positive change after removal can reveal conflict or distraction rather than “negative ownership.”

Coalition credit methods may allocate a chosen utility among sources, but their result inherits the player set, utility, missing-source rule, and generated outputs. PAPER-24 is relevant to stage-specific citation diagnosis, and PAPER-33 is relevant to separating selection from an absorption construct; the W05 course treatment does not elevate either method into a universal influence meter.

9. Attribute failure to the narrowest observable stage

Use the claim–evidence–source graph to diagnose failure. Begin with the answer claim, then inspect citation placement, resolution, passage support, source fit, distinctive-content match, and—only in a controlled context—contribution. A useful taxonomy includes:

  • segmentation failure: one edge is asked to support several separable predicates;
  • resolution failure: anchor or URL cannot identify the source version;
  • placement failure: citation is attached to the wrong claim;
  • entailment failure: passage is topical but does not support the proposition;
  • qualification failure: main value appears but scope or exception is lost;
  • source-fit failure: source says the words but cannot establish the claim type;
  • completeness failure: a material claim has no adequate attribution;
  • absorption uncertainty: content is present but not source-distinctive;
  • wrong-origin failure: a derivative page receives credit for an upstream origin;
  • correctness failure: the answer remains false, stale, or inapplicable despite citation; and
  • use-identification failure: the design cannot distinguish source contribution from redundancy or system change.

Stop at the narrowest supported diagnosis. “The platform ignored the source” is usually too broad. “Under the frozen response, claim C-018 has a visible S-004 edge but the two coders disagree whether the relation is merely cited or unsupported” is auditable. It also identifies the next step: refine placement rules and adjudicate, not rewrite the source for a hidden ranking factor.

10. Annotation agreement is evidence about the instrument

Double coding evaluates whether the annotation instrument is sufficiently explicit for the tested sample. It is not a ritual for producing a high number. Freeze claims and sources; select the subset before labels are compared; keep coders independent; calculate raw agreement by dimension; report class prevalence; preserve all disagreements; and adjudicate with a rule or case rationale.

Cohen’s kappa can summarize chance-adjusted agreement, but it is sensitive to prevalence and small samples. L04 has ten double-coded claims, so its agreement values are descriptive only. Six rows appear in the disagreement log because one claim differs on five dimensions and another differs on absorption. That pattern is more informative than one pooled coefficient: C-018 exposes source-edge and relation ambiguity, while C-020 exposes the boundary between plausible and source-distinctive absorption.

Adjudication should never overwrite the original coder files. Create a new decision record containing claim ID, disputed dimension, both labels, relevant codebook clause, final state or unresolved status, rationale, adjudicator, and date. If a codebook changes, identify which earlier claims require re-review. Recompute denominators after segmentation or eligibility changes.

Automated judges require the same treatment. Measure their sensitivity, specificity, severe-error rate, and stratum shifts against a human reference sample. A correction formula is only defensible if error rates are sufficiently stable for the target population. When models, languages, or claim types differ, show stratum-level results rather than importing one global judge score.

11. Reading the evidence routes without importing their claims

PAPER-24, Diagnosing and Repairing Citation Failures in Generative Engine Optimization, motivates a pipeline-aware failure taxonomy and targeted repair within its reported design. In W05 it supports the question “which observable stage failed?” It does not justify assuming that its improvement magnitude transfers to another corpus, platform, language, or date. Its abstract-level result is not a universal content prescription.

PAPER-33, From Citation Selection to Citation Absorption, motivates measuring more than displayed link counts. Its reported dataset, fetch-conditioned page set, engineered features, and platform snapshot define its ceiling. W05 takes the separation of selection and absorption as a research design idea while treating any textual-influence construct as a proxy, not hidden attention or proven causal dependence.

PAPER-13, Exposing Citation Vulnerabilities in Generative Engines, extends the discussion to publisher attributes, poisoning risk, and the possibility that cited material may not be well reflected in the answer. It supplies bounded adversarial motivation. It does not establish a universal security level for every topic or platform.

PLAT-04, the Bing AI Performance public-preview documentation, illustrates a different evidence type. Official product documentation can define the counts and surfaces it exposes at a checked date. Those counts do not become rank, authority, claim support, or absorption unless the documentation explicitly defines and validates those objects. Platform interfaces are measured surfaces, not peer-reviewed mechanism specifications.

12. A defensible W05 conclusion and exit test

A complete conclusion contains condition, claim unit, source/citation state, measured relation, uncertainty, and boundary. For example:

“In the frozen L04 synthetic fixture, coder A segmented five responses into 25 claims. Fourteen claims carry support edges under the authored codebook; four are contradictions, three are merely cited, and four are explicitly unsupported. Ten claims were double-coded, producing six dimension-level disagreement rows. These labels diagnose the supplied response–source records; they do not establish hidden retrieval, causal source use, public-platform absorption, or general annotation reliability.”

The exit test has five prompts:

  1. Give a case in which a citation is present but the claim is unsupported.
  2. Give a case in which a claim is correct but use of a named source is unproven.
  3. State the denominator for citation entailment and for completeness.
  4. Name two controls in a source-removal comparison.
  5. Rewrite “the source shaped the answer” as one observation and one testable hypothesis.

A learner passes only when every response names an object and evidence route. “The page is cited, so it was used” fails. “The interface displayed S beside claim A; the resolved passage did not entail A; causal use remains untested” passes. The practical goal is not to eliminate uncertainty. It is to place uncertainty at the correct edge of the graph, preserve it, and choose the next test without inventing a hidden mechanism. Precision makes later replication possible.

Frozen L04 Claim–Evidence–Source Graph · synthetic from input to boundary

A five-response synthetic fixture resolves claim spans, source passages, citation edges, support states, six double-coder disagreement rows, and bounded judgments without turning displayed attribution into causal source use.

Controlled source route

Core PAPER-24/PAPER-33 · adversarial extension PAPER-13 · PLAT-04 interface case · Core Notes Chapter 5 · offline L04 claim–evidence–source graph.