# W02 No-Video Transcript — Evidence Before Expression

## Status and use

This is a complete English text equivalent for the planned W02 lecture. **No video has been recorded.** The words below may be read silently, delivered live, or recorded later, but they are not a transcription of an existing class or rehearsal. There are no media timestamps because no media file exists.

At a planning rate of approximately 125–140 spoken words per minute, the narrative is designed to occupy about **35–45 minutes**, excluding optional pauses and discussion. Those values are instructional planning estimates. No classroom, learner, or rehearsal timing has yet been observed. A future recording must be timed from the finished file, captioned, reviewed, and reconciled with this text before replacing the no-video route.

All Northstar Atlas material is synthetic. Paper routes P20 and P40 are discussed through explicit evidence boundaries. Nothing in this transcript claims access to a closed platform’s hidden mechanism or reports an external commercial result.

## Planned chapter budget

| Chapter | Planned spoken emphasis | Text-equivalent function |
|---:|---|---|
| 1–3 | Opening, atomization, claim grammar | Expose the propositions hidden by fluent expression |
| 4–6 | Source identity, ceilings, exact edges | Build the controlled evidence path |
| 7–9 | Dispositions, transport, paper routes | Prevent over-generalization and title transfer |
| 10–11 | Northstar worked resolution | Apply the method to synthetic records and calculations |
| 12 | Release gate and handoff | Produce an inspectable decision and next action |

The chapter budgets are content allocations, not observed durations.

## Chapter 1 — Begin with the sentence you want to believe

I want to begin with one fictional headline. It says: “After Northstar Atlas added FAQ schema, citations across all AI engines doubled and AI-generated sales rose thirty-one percent in six weeks.”

At first hearing, the sentence seems precise. There is a named change, a comparison, a broad system population, a percentage, a business outcome, and a time window. But precision of surface language is not precision of evidence. The sentence compresses several propositions that could fail independently.

First, it says that Northstar added FAQ schema. That is a deployment-state claim. Second, it says that citations doubled. That is a quantitative descriptive claim. Third, the word “after” is used to imply that the deployment caused the change. That is a causal claim. Fourth, the phrase “across all AI engines” transports an observation to a population of systems. Fifth, the sentence says that sales rose by thirty-one percent. That is a commercial measurement claim. Sixth, it calls those sales “AI-generated,” implying that a particular exposure caused them. Those are six different evidence burdens.

A common editorial move is to place one impressive source beside the sentence. Perhaps an official platform document explains structured data. Perhaps a vendor dashboard shows a larger number in February than January. Perhaps a practitioner tutorial says schema helps citations. None of those sources can automatically support the whole sentence. A source that strongly supports the deployment fact can be irrelevant to sales causality. A source that defines one platform feature can be out of scope for other systems.

This is why W02 begins with the atomic claim. The paragraph is too large a unit for evidence review. Citation count is too crude. Source prestige is too general. We need to know which exact proposition is attached to which exact observation, in which source version, under which scope.

Here is the first pause prompt. Take the headline and underline the words *after*, *all*, *doubled*, *generated*, *thirty-one percent*, and *six weeks*. For each word, write the noun or inference it requires. If you cannot name the noun governed by a percentage, the quantitative claim is not yet ready for review.

Evidence before expression does not mean waiting for certainty. It means representing uncertainty before turning it into prose. “Unresolved” is a useful result. “Out of scope” is a useful result. “Do not publish this version” is a useful result. The failure is not admitting what we do not know. The failure is allowing a polished sentence to make the unknown disappear.

## Chapter 2 — Split on independent falsifiability

An atomic claim contains one independently testable predicate. That phrase is a practical editorial rule. It is not a claim that language has one perfect logical decomposition. We split a sentence until each row can receive its own evidence and disposition.

Consider the smaller sentence, “Schema was deployed and citations rose.” It needs at least two rows. The deployment could be true while the citation change is false. The citation change could be present in a dashboard while the deployment record is wrong. If we add “schema caused citations to rise,” we need a third row, because both events could occur while the causal interpretation remains unsupported.

Watch for conjunctions such as *and*, *while*, and *therefore*. Watch for causal verbs such as *caused*, *drove*, *improved*, and *generated*. Watch for comparisons such as *more*, *twice*, and *best*. Watch for quantifiers such as *all*, *most*, and *always*. Watch for platform, population, and time phrases. Also watch for outcome substitution. In this field, mention often becomes citation, citation becomes visibility, visibility becomes traffic, and traffic becomes revenue, all inside one paragraph. Those events may be related, but they have different units and denominators.

Atomicity lets us preserve useful facts when an adjacent claim fails. Suppose an internal manifest records that 118 pages received a build. That row can remain supported even if the claim of citation impact is unresolved. Suppose a monitor records 24 events in one period and 41 in another. We may report those exported values with caveats even if we cannot attribute the change to schema. Suppose a CRM counts assisted transactions. We may describe the dashboard rule without calling those transactions incremental sales.

Do not atomize by merely cutting a sentence at every comma. Find independent truth conditions. “In the frozen monitor, visible citation incidence was 24 of 100 opportunities in January and 41 of 100 in February” can remain one comparative row if both values, denominators, windows, and protocol belong to the same estimand. “The increase occurred because of FAQ schema” is another row because it asks for a counterfactual.

Preserve the original expression separately. The original text and its position matter because later reviewers need to see what changed. The ledger should include the original span, the atomic rewrite, and the release action. If the rewrite overwrites the source sentence, the audit loses the very object it is supposed to govern.

Pause again. Take a compound claim from your preparation. Circle each independently falsifiable predicate. If one row could be false while the rest remained true, split it. Then label each row descriptive, comparative, causal, normative, predictive, or prescriptive. That modality determines what kind of evidence bridge you need.

## Chapter 3 — Give the claim a grammar and a boundary

An atomic row needs more than short wording. It needs a structured boundary. Begin with a stable claim ID. Record the subject, predicate, and object or value. Then name the population or unit, system and surface, time window, locale, conditions, modality, and uncertainty.

The subject might be a page set, a source, a platform, a response, a user group, or a metric. The predicate might assert a state, comparison, cause, requirement, or forecast. The object might be a numerical value, category, or other entity. The population tells us whether the claim concerns prompts, responses, pages, sessions, users, or organizations. The system field prevents one API or interface from becoming an unnamed universal engine. The time field prevents an old snapshot from becoming a present-tense property.

The conditions field records what must be true for the claim to apply: account state, locale, prompt set, candidate corpus, model snapshot, run count, or evaluation protocol. The modality field tells us whether an observation is being described or a cause is being asserted. The uncertainty field contains intervals, missingness, disagreement, and unknowns. Unknown is a value. It is not a blank that an editor may fill by intuition.

Let us strengthen “citations increased.” A more inspectable descriptive claim is: “In the synthetic Northstar monitor, on the named Aster Answer surface, for the en-US fifty-prompt set and two runs per prompt, the export labeled twenty-four of one hundred January responses and forty-one of one hundred February responses as containing at least one visible citation to a Northstar help page.” That wording names the measurement object. It still does not tell us that the parser was perfect or that the prompts were unchanged. Those conditions belong in the edge and caveat.

Now notice what happens when we change the verb. “The export labeled” is a statement about a file. “Visible citations increased” is a statement about measured interface events. “Schema increased citations” is a causal statement about an intervention effect. Each move requires new evidence. Correct grammar makes that movement visible.

Quantifiers create another boundary. “All AI engines” requires a defined population. What counts as an engine? Are APIs and consumer interfaces separate surfaces? Which locales and dates are included? Observing one fictional surface does not partially support the word *all*. The population claim is simply outside the measurement scope.

Your next check is simple. Read each atomic row and ask: could a second reviewer identify the unit, surface, and time without looking elsewhere? Can the reviewer tell whether the verb is descriptive or causal? If not, the row is short, but it is not yet atomic enough for evidence governance.

## Chapter 4 — Identify the source before extracting authority

Now we turn from claims to sources. A source is not just a link. Record the issuer, authors, type, title, version, publication status, date, stable identifier, canonical location, method, artifacts, license, sponsor, correction path, and unknown fields.

These fields matter because different documents change in different ways. Platform documentation can be revised without a scholarly version history. A preprint can receive a new version. A repository can change defaults between commits. A video can preserve a demonstration of an interface that no longer exists. A vendor report may rely on a proprietary parser and sample. A standard can be superseded. When the version changes, affected claim rows may need review.

Identity cues do not complete the evidence record. A DOI helps identify a paper. It does not validate the sampling frame or causal design. An institutional logo identifies an issuer. It does not prove that an organization implemented the stated policy. A GitHub repository exposes code. It does not establish that the result transports to another dataset or production setting.

Independence is part of source identity. Imagine nine source cards: a release note, deployment manifest, internal change log, methods memo, monitor export, parser QA note, CRM report, platform help card, and practitioner tutorial. Nine documents do not necessarily represent nine independent bodies of evidence. Four internal operations documents may share one origin. The monitor and parser note belong to one instrumentation family. A partner article that repeats a vendor report belongs to the report’s family unless it conducted new work.

Dependence does not make a source useless. Multiple internal records can triangulate version, deployment, and process. But counting them as independent confirmation would exaggerate the evidence. The ledger should record common origin, syndication, shared dataset, overlapping authorship, and common sponsor when those relationships matter to the claim.

Here is the source-identity check. Choose one source you consider highly credible. Complete this pair: “This source can support…” and “This source cannot establish…”. An official help page may support a documented feature’s syntax and scope. The same page cannot establish that Northstar’s pages were processed, that citations changed because of the feature, or that another provider behaves the same way. Authority is relative to the claim.

## Chapter 5 — Use ceilings as roles, not scores

The course catalog uses seven evidence-ceiling codes: L, N, A, B, C, D, and E. They are arranged as roles, not a prestige ladder.

L means a local controlled artifact. A hash manifest can establish byte identity and file state. It cannot make the external assertions inside the file true. N means a normative or official authority within its stated scope. A standard can establish a requirement or definition. It cannot by itself establish compliance, implementation, or performance effect.

A means primary research bounded by design, data, models, date, and uncertainty. It can support a result under those conditions. It cannot automatically describe every current commercial system. B means academic pedagogy or a maintained research implementation. A course can explain a method, and a repository can provide a reproducible route. Neither replaces the primary evidence for an empirical conclusion.

C means first-party platform documentation for the named platform, surface, locale, and date. It can establish documented behavior, eligibility, or policy. It cannot disclose all hidden mechanism details or prove cross-platform effects. D means vendor observational or quasi-experimental evidence. Such work may contain useful large operational samples, but causal and market-wide language depends on design transparency. E means practitioner or secondary synthesis. It can supply vocabulary, workflow examples, hypotheses, and contemporary questions. It requires corroboration for scientific, legal, standards, and causal claims.

Do not average these letters into a credibility score. Scoring invites compensation. A famous issuer may hide an absent denominator. Many dependent sources may hide one origin. A polished visual may hide an undefined outcome. Instead, assign the ceiling relative to the proposed claim and write the positive and negative boundary.

This also explains why a source can be credible yet unsuitable. A platform document can be the strongest available source about its own declared schema eligibility and still be irrelevant to a sales-effect claim. A randomized study in one environment can be excellent research and still be out of scope for a different system version. A practitioner tutorial can be useful for discovering an implementation hypothesis and still be unable to carry scientific or causal language.

For your pause prompt, select four source types: one paper, one official document, one vendor report, and one tutorial. Give each a one-sentence admissible use and a one-sentence forbidden extrapolation. If your boundary is “this source may be biased,” make it more precise. Name the claim class, population, design, or system that the source does not cover.

## Chapter 6 — Build an exact Claim–Evidence–Source path

The evidence path has three layers. At the top is the atomic claim. In the middle is the exact evidence item: a passage, table cell, data slice, record, or observation. At the bottom is the versioned source identity that contains or governs that item.

The middle layer prevents a whole document from becoming a universal support token. A reviewer needs a page, section, paragraph, table, figure, record ID, line range, or verified recording timestamp. Because no W02 recording exists, this package has no timestamps. In a future video review, a timestamp would locate the speaker’s claim; it would not make the claim true.

Every edge has a type. It may support, partially support, contradict, qualify, be non-informative, or be out of scope. Topical similarity is not entailment. A blog can contain the words *FAQ schema*, *citations*, and *leads* while providing no evidence for a schema-to-sales causal effect.

Carry three boundaries with each edge. The claim boundary states the exact proposition. The evidence boundary states what the passage or observation shows. The source boundary states what the source’s method, authority, version, and independence permit.

Suppose a deployment manifest says 120 pages were targeted, 118 completed, and 116 passed markup validation. That exact record contradicts “successfully deployed to all 120 pages” but partially supports a narrower deployment statement. It says nothing about whether a public system fetched, indexed, retrieved, or cited the pages. The edge is strong and local.

Suppose a platform help card says supported structured data can assist interpretation but does not ensure retrieval or visible attribution. That passage supports the platform’s declared boundary. It does not support an observed outcome for Northstar. A source can be directly relevant to the audit because it limits a claim, not because it provides a positive result.

Your check is reconstructability. Could a second reviewer open the exact source version, find the cited location, see the relation type, and understand the scope match? If the ledger only says “platform documentation” or contains a homepage URL, the path is not resolvable.

## Chapter 7 — Preserve five dispositions and five transport tests

W02 uses five primary dispositions. Supported means the evidence entails the claim within compatible scope and no unresolved critical conflict is known. Partially supported means a narrower proposition survives, but a value, quantifier, modality, or scope must change. Contradicted means comparable evidence supports an incompatible proposition. Unresolved means the record conflicts or cannot distinguish important explanations. Out of scope means the evidence never observed or governed the named population, system, time, jurisdiction, or outcome.

“No evidence found” is not the same as “false.” It records the search boundary. Likewise, contradiction should not be deleted. Compare scope, denominator, method, version, date, and source dependence. Sometimes two values are not contradictory because they use different units. Sometimes the newer source supersedes the old one. Sometimes both are credible but sample different populations. If the comparison cannot resolve the difference, preserve both edges and remain unresolved.

Then apply five transport tests. Population transport moves from a sample to a broader group. System transport moves from one model, API, or interface to others. Time transport moves from a dated snapshot to a present-tense property. Causal transport moves from association or before-and-after change to an intervention effect. Outcome transport moves from one event to another, such as citation to referral or assisted transaction to incremental sales.

For each move, name the bridge. A population bridge may require a sampling and transport argument. A system bridge may require a defined multi-platform protocol and heterogeneity analysis. A causal bridge requires a suitable comparator and treatment design. An outcome bridge requires linked units and a valid construct.

Percentages receive a separate check. Every percentage needs a noun and denominator. If exported visible-citation incidence changes from 24 of 100 response opportunities to 41 of 100, the absolute change is 17 percentage points. The relative change is 17 divided by 24, about 70.8 percent. Both calculations are correct. Neither is a doubling. Doubling 24 under the same denominator would require 48.

Correct arithmetic does not repair the instrument. If four prompts changed and the parser has uncorrected errors, the export values remain bounded observations. They are not a clean causal effect. The disciplined sentence names the export, unit, windows, and measurement limits.

## Chapter 8 — P20: do not turn model choice into source truth

The first core paper route is P20, *What Evidence Do Language Models Find Convincing?* The reviewed study constructs controversial binary questions and pairs passages supporting opposing conclusions. Several language models answer from those passages. The analysis asks which passage-level features are associated with the stance the model adopts and how model-generated edits change that choice.

This is useful for W02 because it separates passage influence from source authority. In the controlled setup, semantic relevance is the most consistent positive signal among the examined features for most displayed models, while several style features show weak, mixed, or negative results. The study also reminds us that a model may favor one passage without establishing that the passage is true.

Now state the boundary. The passages are decontextualized. Source URLs, publisher identity, dates, headings, visuals, and author credentials are removed. The task uses two passages and yes-or-no answers. Questions and filtering are constructed, effective samples differ by downstream model, and one displayed model cell is especially small. The result does not measure open-web retrieval, visible citation, current multi-source production behavior, source credibility, factual truth, or human persuasion.

The word *convincing* in the title must not be transferred into a stronger claim. In this route, “winning” means that the generated answer adopts a passage’s stance under a controlled conflict protocol. It does not mean credible, correct, independent, or ethically desirable. Relevance to a query is not an evidence ceiling.

P20 therefore supports a bounded lesson: model decisions can be sensitive to passage content and semantic relation within a specified setup. It cannot serve as a universal writing recipe or proof that a relevant passage will be retrieved, cited, believed, or rewarded across systems.

Pause and complete the sentence: “P20 can support a discussion of…” Then complete: “P20 cannot establish…” If either sentence uses the words *all models*, *production engines*, *truth*, or *human behavior* without qualification, revise it.

## Chapter 9 — P40: more evidence can still be a weak graph

The second core paper route is P40, *GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning*. The reviewed benchmark gives verification procedures a fixed set of three evidence objects. It progressively replaces zero, one, two, or three of those objects with one of several synthetic or tampered attack types. It compares multiple verification procedures and reports accuracy and efficiency measures.

This route is useful because it challenges two shortcuts. First, clean performance does not settle robustness when evidence changes. Second, no method in the reviewed results dominates across attack families, replacement ratios, categories, and model backbones. A procedure that resists an instruction-like attack may remain vulnerable to content tampering. Evidence aggregation must therefore preserve attack and source-family structure rather than collapse everything into one confidence score.

The boundary is equally important. Each fixed configuration is evaluated once over the benchmark. The reviewed version does not provide confidence intervals or repeated-seed sensitivity. Important implementation, model snapshot, prompt, annotation, and release details are incomplete. The benchmark selects exactly three evidence objects per claim and replaces exact counts. That controlled design does not estimate how often real-world sources are compromised, how attackers control live retrieval pools, or which method is universally best.

P40 does not prove that a source set in this course has been poisoned. It does not establish production prevalence or user harm. It does not turn residual accuracy under full replacement into successful evidence use, because an answer may come from parametric knowledge or other behavior.

For W02, the safe lesson is structural. Source count is not independence. Agreement is not validity when sources share an origin or manipulation. Robustness is attack-specific. A claim ledger should preserve evidence families, contradictions, uncertain provenance, and the possibility that an apparently corroborating set is not genuinely independent.

Here is the paper-route check. Explain why P40 is relevant to evidence governance without repeating any benchmark point estimate. If you can state the lesson accurately without a headline number, you are less likely to treat the number as a portable fact about production systems.

## Chapter 10 — Resolve the Northstar deployment and citation claims

We now return to the synthetic Northstar case. The first source card is a change-request release note. It targets 120 English help pages and lists a bundle: FAQ JSON-LD, revised headings, updated answer summaries, and related-article links. This supports the planned target and bundle. It does not prove successful release.

The deployment manifest records 120 targeted pages, 118 completed, two rolled back, and 116 with valid FAQ markup. Our first atomic claim says that FAQ schema was successfully deployed to all 120 pages on 3 February. That claim is contradicted as written. A narrower statement is partially supported: 118 pages completed the build and 116 passed the recorded markup validation. Notice how a strong local artifact preserves a useful fact while rejecting the word *all*.

The monitor export covers one fictional surface, Aster Answer, in en-US. It contains fifty prompt IDs and two runs per prompt in each monthly window. The export labels 24 of 100 January responses and 41 of 100 February responses as containing at least one visible Northstar citation. Four prompt wordings changed. A parser QA sample contains two false positives and one false negative, and the interface changed during February.

Our second claim says citations doubled. It is contradicted by the supplied values. The exported incidence changed from 24 percent to 41 percent: 17 percentage points, or about 70.8 percent relative to the January value. The parser and prompt limits still apply. We should not repair them with arithmetic.

The third claim says FAQ schema caused the increase. The site change log records revised headings, answer summaries, internal links, lower latency, and publicity in the same period. The methods memo records no concurrent control, random assignment, interrupted-series design, or drift block. Four prompt wordings and the monitored interface changed. The causal claim is unresolved and unsupported for publication. Sequence is not enough.

The fourth claim says the result occurred across all AI engines. The methods memo explicitly says only Aster Answer was monitored. That claim is out of scope. It is not partially supported by a one-platform result. To investigate transport, we would first define the population of platforms and surfaces, freeze configurations, and report heterogeneity rather than replacing failures with a pooled slogan.

Pause here and ask which source is credible yet unsuitable. The fictional Aster help card is authoritative about Aster’s declared structured-data boundary. It says eligibility does not ensure retrieval, inclusion, or visible attribution. That source is useful because it limits the causal claim. It is unsuitable as evidence that Northstar obtained an effect.

## Chapter 11 — Resolve the commercial outcome and rewrite the expression

The CRM export supplies two counts. January has 260 sessions with an AI-referral tag and 13 transactions with any such session in the prior 30 days. February has 390 tagged sessions and 17 assisted transactions. The assisted-transaction count increased by four. Four divided by thirteen is approximately 30.8 percent.

That calculation supports a bounded descriptive claim: under the rule named `assist-30d-v1`, the dashboard count increased from 13 to 17. It does not support “AI-generated sales.” The report contains no revenue or margin. The rule counts a transaction as assisted if an earlier tagged session occurred in the attribution window. Direct visits can still receive the label. The CRM and citation monitor have no common response, session, page, or user key.

Our fifth atomic claim—the assisted count changed by approximately thirty-one percent under that rule—is supported as a dashboard description. It needs the raw counts and operational definition. Our sixth claim—schema caused thirty-one percent more sales—is unsupported. It has a construct mismatch, because assisted transactions are not incremental sales, and a causal mismatch, because the records are not linked and no counterfactual design exists.

Now rewrite the headline. A bounded internal summary might say: “In this synthetic exploratory monitor, exported visible-citation incidence on Aster Answer changed from 24 of 100 recorded opportunities in January to 41 of 100 in February. Prompt wording, interface, parser context, page content, links, latency, and publicity also changed, so the record does not identify an effect of FAQ schema or support transport to other surfaces. Separately, the fictional CRM counted 13 and 17 transactions with an AI-referral-tagged session in the prior 30 days; that operational metric is not incremental sales and is not joined to the citation monitor.”

The rewrite is longer because it restores the objects the headline collapsed. It names the surface, unit, denominators, windows, co-interventions, causal boundary, transport boundary, and separate CRM construct. A shorter external sentence could say that an exploratory monitor showed a change requiring controlled follow-up. It should not name a schema effect, all engines, or a sales effect.

This is not weak communication. It is communication that another analyst can reconstruct. The evidence has not disappeared into tone. The strongest surviving facts remain available, and the unsupported propositions become explicit research questions.

## Chapter 12 — Release the claim, not the uncertainty

We finish with a release gate. For every atomic row, choose one action: approve as written, approve after narrowing, publish with a visible caveat, seek specified evidence, escalate for specialist review, or block. Do not force all rows into supported or rejected. A good ledger preserves a mixture.

Before approval, ask whether a second reviewer can reconstruct five things. First, what is the exact atomic claim? Second, which evidence locator and source version bear on it? Third, what ceiling and edge type apply? Fourth, what contradiction, dependence, missingness, or unknown remains? Fifth, what release action and next evidence step follow?

The next evidence action should be proportionate. To close the deployment gap, validate and document the rolled-back pages. That establishes technical state, not effect. To improve the parser estimate, double-code a larger stratified sample across interface versions and report disagreement. That improves measurement, not causality. To estimate a schema effect, use a preregistered intervention with a suitable control, blocked repetitions, and locked co-interventions. That may estimate a local effect under the design, not a timeless cross-engine law.

To study transport, define the platform and surface population, sample named configurations, and preserve heterogeneity. To study a commercial effect, first define the outcome, create a lawful and privacy-appropriate link between exposure and outcome units, and choose a design that addresses confounding and interference. A larger dashboard does not substitute for those steps.

Your exit ticket has two parts. First, name one source that is credible yet unsuitable for a claim people commonly attach to it. State exactly why: wrong system, wrong population, wrong time, wrong modality, wrong outcome, or insufficient design. Second, write one atomic claim with a source version, exact locator, ceiling, disposition, release action, and cheapest ethical next evidence step.

The central lesson is simple. A document is not an evidence edge. Authority is claim-relative. Repetition is not independence. Sequence is not cause. Citation is not sales. When the evidence cannot carry the expression, revise the expression—not the evidence boundary.

That completes the no-video route for W02. The next practical step is L04’s seed Claim–Evidence–Source ledger, where these rows become a machine-readable graph and disagreements remain visible for adjudication.
