# W10 Lecture Notes — Entities, Authority, and Source Ecosystems

## 1. Identity must be resolved before authority is judged

A fact is not ready for corroboration until its subject is stable. Names are labels, not identities. The same organization can publish under a legal name, a trading name, an abbreviation, a translated name, and a former name. The same abbreviation can belong to several organizations. A product can inherit a parent-company name, split into editions, or retain an old label after ownership changes. A search result that looks relevant can therefore describe the wrong entity while matching every visible word in a query.

Begin with an entity dossier rather than a relevance judgment. Its minimum header is a course-controlled stable ID, canonical name, entity type, jurisdiction, recognized aliases, rejected aliases, official or registry identifiers, effective dates, and resolution status. The stable ID is not a metaphysical truth. It is a durable key within the governed project, connected to external identifiers and an explanation of how candidates were accepted or rejected.

Jurisdiction is often decisive. A registry number can be unique within one legal system yet collide elsewhere. A brand can be licensed to different operators by country. A policy may apply to an EU service but not a similarly named US service. Record jurisdiction as a field, not as an assumption in prose. If the evidence does not resolve the candidate, retain `ambiguous` rather than selecting the most familiar name.

The Northbridge case uses `ENT-NBDC-001`, canonical name `Northbridge Data Cooperative Ltd.`, fictional jurisdiction `EX-NORTH`, and registry key `EXN-88421`. `Northbridge Data` and the historical `Northbridge Cooperative` are accepted aliases with dates. `NDC` is rejected as a stand-alone resolver because it maps to multiple supplied candidates. The decision becomes reproducible because another reviewer can apply the same keys.

**Teaching check:** if a learner says that a well-known publisher settles identity, ask which identifier connects the page to the candidate. Publisher reputation cannot repair an unresolved subject.

## 2. A claim needs scope, time, and status

After identity comes claim construction. “Northbridge retains exports for 90 days” is still too loose. A reconstructable claim records an entity ID, predicate, value and unit, product or plan, effective interval, locale or jurisdiction, qualifying conditions, and status. In the fixture, `CLM-RETENTION-001` states that `ENT-NBDC-001` standard service retains audit exports for 90 days, for the `EX-NORTH` service, effective from 15 July 2026, status `current-under-review`.

Atomization matters. A sentence that says a service “stores exports securely for 90 days and meets all regulations” bundles retention duration, security, and legal compliance. Those predicates require different evidence. Split them before evaluating support. A source can support one edge while contradicting or ignoring another. The governed representation should preserve that difference.

Time is not a decorative citation date. A source may have a publication date, observation date, effective date, reviewed date, retrieved date, and superseded date. These answer different questions. A current page can quote an old policy. An undated page can be retrieved today yet remain stale. A recent aggregator can repeat a superseded original. Use explicit fields and never infer freshness solely from the crawl or access date.

Status also controls language. `proposed`, `current`, `superseded`, `disputed`, `corrected`, and `unresolved` should not collapse into a binary true/false label. A claim may be current for the standard plan and superseded for the legacy plan. Another may be supported by an owned primary record but await independent corroboration. The final prose must reflect that state.

PAPER-01 contributes only a time-provenance failure lesson. Its local full text contains collection or API dates after the 24 August 2026 course freeze. That mismatch triggers quarantine. W10 takes no substantive evidence from PAPER-01; it uses the record only to show that a plausible title, DOI, and PDF cannot override incoherent time provenance.

## 3. Authority is claim-relative

Authority is not a permanent score attached to a domain. It is a relationship among a source, a scoped claim, a time, and an evidentiary role. An official corporate registry can be highly authoritative for legal name, registration number, and jurisdiction. It is usually irrelevant to a product feature. A product manual may be the primary record for the vendor’s current specification. It is not automatically independent evidence that the product outperforms competitors. A regulator may be authoritative for the text and status of its own order, while a later court record may control the disposition of an appeal.

Use an authority matrix. Put atomic claims in rows and candidate sources in columns. For each cell, record applicability, directness, temporal fit, jurisdictional fit, independence, and known dependence. The matrix prevents a source that is strong for one row from lending unearned authority to every other row.

This is why “owned” is not synonymous with “weak.” If an organization defines its current price, API limit, plan name, or correction policy, its authorized documentation may be the primary source. The weakness appears when the claim requires external assessment: “most trusted,” “more reliable than competitors,” or “widely recommended” cannot be established merely because the organization says so. An independent source can be valuable for those claims, but independence alone does not guarantee expertise, accuracy, freshness, or relevance.

PAPER-14 provides a metadata-limited research-context route for reputation and discovery. The latest local CSV audit classifies it as Notes catalog-only and reports no URL, abstract, or access date in the export; the local catalog also marks it as a preprint whose venue is not confirmed here. Catalog routing is not substantive absorption. Use PAPER-14 only to frame questions, not to grant a universal authority score or support the case facts. PAPER-35 supplies a core route for evidence-ecosystem and trajectory thinking. Its particular architecture and evaluation remain study-specific. Neither paper validates the fictional Northbridge facts.

**Decision rule:** describe what a source is authoritative for, not whether it is “authoritative” in the abstract.

## 4. Primary, owned, and independent are separate axes

Primary, owned, and independent answer different questions. A primary source directly records or produces the fact at issue: a signed contract for its terms, an official registry for a registration, a product owner’s current specification for its declared setting, or a researcher’s dataset for what that researcher collected. An owned source is controlled by the entity whose interests are implicated. An independent source is organizationally and materially separate from the claim origin.

These properties can overlap. A current vendor document can be both primary and owned. An external audit commissioned by the vendor can include direct observations but have financial dependence that must be disclosed. A news story can be institutionally independent yet materially dependent if it merely rewrites a press release. A database can look third-party while ingesting the same vendor feed. Classify the relation, not the page design.

The Northbridge current documentation is primary and owned for the 90-day declared specification. The fictional procurement audit is independent and directly tests export availability through day 90 under the named plan. The registry is independent of the vendor’s publishing channel and primary for legal identity, but it is inapplicable to retention. The syndicated story is published elsewhere but depends on the owned press release. Its external domain does not create independent corroboration.

For each source, record origin organization, author or responsible unit, funding or sponsorship when known, directness, method, quoted or transformed source, publication and effective dates, correction route, and claim applicability. “Third party” is too vague to substitute for those fields.

PLAT-13 documents a shared structured-data vocabulary. It may help represent an entity, alias, identifier, or page relationship consistently. It does not verify the represented values, establish publisher independence, or demonstrate a visibility effect. Syntax and semantics make a record machine-readable; governance and evidence make it accountable.

## 5. Provenance graphs expose dependence

A source list shows items. A provenance graph shows relationships among them. Use nodes for entity records, atomic claims, source artifacts, origin organizations, and corrections. Use typed directed edges such as `asserts`, `supports`, `observes`, `quotes`, `syndicates`, `transforms`, `contradicts`, `supersedes`, and `corrects`. Every edge carries a source ID, date, scope, and confidence or review state.

Dependence is about material reliance, not only ownership. Suppose an owned press release states 90 days. A wire copy republishes it, a news page lightly edits the wire copy, and an aggregator extracts the news sentence. Four URLs now display the value. They still form one origin cluster if none performed independent verification. Count four publication nodes, one claim origin, and zero added independent corroborators from that branch.

The graph should not hide distribution. Syndication can matter for reach, correction burden, and staleness. It simply answers a different question from corroboration. Record both publication count and independent-origin count. A robust report might say: “The claim appears on four supplied pages, three of which depend on one owned origin; one separate audit supplies independent direct observation.”

Original visual design for this lesson uses five node shapes: rounded rectangle for entity, hexagon for claim, document for artifact, circle for origin organization, and diamond for correction event. Solid arrows mean direct assertion or observation; dashed arrows mean syndication or transformation; double arrows mean conflict; dotted arrows mean requested refresh. Labels duplicate every pattern, so color is never required.

A table equivalent lists `from_id`, `to_id`, `edge_type`, `origin_cluster`, `effective_date`, `review_status`, and `note`. Sorting first by claim and then by origin cluster lets a learner reproduce the count without seeing the graph.

## 6. Corroboration requires applicable independent origins

Corroboration is stronger than repetition. For this package, a claim is corroborated within the frozen case when at least one applicable primary record and at least one materially independent, applicable source origin converge on the same scoped value, with conflicts preserved and assessed. This is an instructional rule, not a universal scientific threshold. High-risk claims may require more origins, domain expertise, legal review, or direct testing.

Applicability comes first. The registry cannot corroborate a retention duration because it does not address the predicate. A current owned document supports the vendor’s declared specification. The independent audit supports observed availability through day 90 under the matching plan and time. Together they support a bounded statement. The syndicated and aggregator pages add reach but not independent origins.

Avoid converting corroboration into a vote. Two copied pages do not outvote one direct record. A later source does not automatically defeat an earlier source if the later page describes a legacy plan. A prestigious outlet does not repair a missing method. Evaluate edges and scopes.

The bounded Northbridge conclusion is: within the supplied synthetic records, as of 20 July 2026, the standard-plan 90-day retention claim has current owned primary support and one applicable independent audit origin. It is not a conclusion about any real company, hidden retrieval system, ranking, citation, traffic, or business outcome. The registry resolves identity only; dependent copies do not increase the corroboration count.

Ask learners to report three counts: pages displaying the claim, materially distinct origin clusters, and applicable independent corroborating origins. This small discipline exposes false abundance.

The count also needs an uncertainty state. Dependence is sometimes explicit through a syndication label or quotation, sometimes inferred from a documented feed, and sometimes unresolved. Use `dependent-confirmed`, `dependent-probable`, `independence-confirmed`, and `unknown` only when the project defines evidence for each state. Do not award independence merely because dependence was not discovered. In the frozen fixture, the news and aggregator edges are explicitly supplied, so clustering is deterministic. In field work, an unknown relation should appear in sensitivity reporting: present the origin count under the narrow confirmed-dependence rule and under a conservative possible-dependence rule. This does not turn the graph into a probability model. It shows which conclusion depends on an unresolved provenance edge and directs the next verification task.

## 7. Conflicts are records, not debris

Contradictory sources should remain in the ledger. Deleting an old or inconvenient record removes the evidence needed to understand change, detect propagation, and audit a correction. Instead, classify the conflict.

Useful states include: exact contradiction under matching scope; temporal supersession; plan or population mismatch; jurisdiction mismatch; definition mismatch; transcription error; retraction or correction; unverifiable attribution; and unresolved conflict. Each entry records both claim versions, entity key, source IDs, dates, affected scope, provisional decision, reviewer, and next action.

The Northbridge fixture includes an older owned PDF saying 60 days. Its effective scope predates the 15 July 2026 update, so it is marked `superseded` rather than erased. An independent comparison page says 30 days, but it is dated February 2025 and describes the legacy plan. That is a scope and time mismatch, not a current-plan refutation. Its record remains visible because downstream pages may still repeat it. If a source said 30 days for the current standard plan after 15 July 2026 and supplied a credible method, the conflict would remain unresolved until investigated.

Staleness is also claim-relative. A five-year-old incorporation record can remain current; a two-week-old price page can already be obsolete. Define review intervals and triggers by predicate risk and volatility. Do not equate “old” with “wrong” or “recent” with “current.”

PAPER-01 illustrates a different conflict: internal collection or API dates extend beyond the course freeze. The package quarantines the record instead of inventing a reconciliation. This is provenance discipline, not a substantive judgment about the paper’s claims.

## 8. Correction must be authorized and traceable

A correction loop begins with a disputed atomic claim, not a campaign for more mentions. The responsible team preserves the current record, verifies entity and scope, locates the authorized primary source, assigns an owner and independent reviewer, and decides whether the claim is wrong, stale, ambiguous, or unsupported. If correction is warranted, the owner changes the source they are authorized to maintain and records the prior value, new value, reason, approver, effective date, and rollback target.

Downstream publishers are approached through legitimate correction channels: an errata form, editorial contact, database feedback route, or contractually authorized feed. The request cites the corrected primary record and the precise affected statement. It does not demand favorable coverage, fabricate independent articles, conceal sponsorship, or create mass copies. Each downstream node receives a state such as `notified`, `acknowledged`, `corrected`, `declined`, `unresponsive`, or `not-contacted-no-authority`.

Correction propagation is not instantaneous and cannot be assumed. A publisher may preserve an old article as historically accurate while adding a note. An aggregator may refresh on a schedule. A cached copy may remain. The ecosystem graph records the state of each branch rather than claiming that the “web is fixed.”

Separate evidence correction from system outcome. Improving an owned document can make the primary record clearer and more current. That does not prove that a crawler retrieved it, a model used it, an answer cited it, or a user acted on it. Those require separate measurements.

## 9. Refresh triggers, review dates, and rollback

Calendar review is necessary but insufficient. Define event-based triggers: change of legal name or jurisdiction; new registry identifier; ownership transfer; product or plan revision; policy effective date; source correction or retraction; alias collision; newly discovered independent contradiction; broken correction link; expired audit; and missed review deadline. Every trigger opens a dated review task with an owner, scope, and disposition.

The graph itself is versioned. A frozen reproduction package contains the entity dossier, claim table, source cards, edge table, origin-cluster assignments, conflict ledger, correction log, and file hashes. A deterministic build sorts records by stable ID, validates referential integrity, rejects unknown edge types, recomputes origin counts, and emits the bounded conclusion from explicit rules. Another reviewer can reproduce the result without network access.

Rollback is not merely retaining an old file. Define the exact artifact pair or dataset version, restoration command, content hashes, authorization, test, and post-rollback review. LAB-L05 demonstrates this at local-fixture scale. The control and treatment pages contain identical locked claim text and source identities, while evidence-layout proximity is the sole declared factor. The audit writes a rollback manifest connecting the treatment pair to the control pair and checks factual equivalence.

A LAB-L05 pass means three locked claims, three source identities, one changed factor, and zero audit errors under the frozen code and data. It does not show human comprehension improvement or any external crawler, retrieval, ranking, answer, citation, traffic, or business effect. The experimenter and checklist reviewer remain distinct synthetic roles.

## 10. Interventions should improve evidence, not manufacture independence

Legitimate intervention categories include clarifying an authorized primary record, separating atomic claims, adding visible dates and scope, publishing a change log, linking evidence spans, exposing correction channels, improving accessible structure, and notifying real downstream publishers of verifiable errors. These actions make evidence easier to inspect and govern.

Prohibited shortcuts include invented third-party publication, undisclosed sponsorship, purchased coverage represented as independent, mass duplication, circular citation networks, deceptive attribution, hidden content variants, and fabricated statistics. They corrupt the independence model and can cause harm even if they increase page count. The studio stops when a proposal depends on any of these tactics.

PAPER-14 can motivate discussion of reputation environments, and PAPER-35 can motivate evidence-ecosystem representations, but W10 does not extract a recipe for manipulating a system. PLAT-13 may help standardize a visible entity representation, while its evidence ceiling stays at vocabulary semantics. PAPER-01 remains audit-only. Source status and boundaries travel with every reference.

An intervention brief should state authorized surface, target claim, changed factor, preserved factors, expected human or governance mechanism, measurement plan, stop conditions, reviewer, version, and rollback. If the intended effect is clearer human review, do not relabel it as a ranking intervention. If no system response is measured, say so.

## 11. Bounded reporting and exit test

A strong final statement has five parts. First, identify the entity and resolution keys. Second, state the atomic claim with scope and effective date. Third, name applicable support by role. Fourth, report dependency clusters and conflicts. Fifth, state the evidence ceiling and unresolved work.

For the fixture: `ENT-NBDC-001` is resolved to Northbridge Data Cooperative Ltd. in `EX-NORTH` using registry key `EXN-88421`; `NDC` alone is insufficient. `CLM-RETENTION-001` concerns the standard service, 90 days, effective 15 July 2026. The current owned documentation is primary for the declared specification, and one independent audit origin reports matching direct observation through day 90. Two downstream pages belong to the owned-origin syndication cluster and add no independent corroboration. The 60-day PDF is superseded; the 30-day comparison is legacy-plan evidence. The conclusion is frozen as of 20 July 2026 and requires refresh on a plan, source, registry, or conflict trigger.

This package does not establish behavior of a real entity or platform. It does not infer retrieval, answer use, citation, ranking, traffic, reputation change, or business outcome. A source graph is an accountable evidence representation, not a causal model of hidden systems.

Final self-check:

1. Could another reviewer resolve the same entity from my recorded keys?
2. Does each claim have one predicate, explicit scope, dates, and status?
3. Did I judge authority per claim rather than per domain?
4. Did I count independent origins instead of URLs?
5. Are conflicts, stale states, and nonresponsive branches preserved?
6. Is every correction authorized, reviewable, and reversible?
7. Did I keep PAPER-01 inside its time-provenance quarantine?
8. Did I state what PAPER-14, PAPER-35, PLAT-13, and LAB-L05 cannot establish?

If any answer is no, the ecosystem is not yet ready for a corroboration claim.
