# W02 Lecture Notes — Evidence Before Expression

## 1. The opening claim

Consider a fictional vendor headline:

> After Northstar Atlas added FAQ schema, citations across all AI engines doubled and AI-generated sales rose 31% in six weeks.

The sentence feels economical. It names an intervention, an apparent change, a broad population, a commercial outcome, and a time window. That economy is precisely the audit problem. At least six propositions are compressed into one line: a specific change occurred; observed citations changed; the change was caused by FAQ schema; the result held across all engines; an outcome called “AI-generated sales” changed; and the content intervention caused that outcome. Each proposition needs a different object, denominator, comparator, and evidence route.

A single strong-looking citation beside this sentence can launder the other propositions. A deployment log might establish that schema was added, but not that citations increased. A dashboard might establish a descriptive change in one monitored surface, but not causation. A customer-relationship report might count assisted transactions, but not sales caused by an answer engine. W02 begins from a strict rule: **the unit of audit is the atomic claim, not the paragraph, document, or citation count**.

Evidence before expression does not mean that writers must wait for perfect knowledge. It means that uncertainty is structured before prose is optimized. A defensible record can say supported, partially supported, contradicted, unresolved, or out of scope. The failure is not uncertainty. The failure is turning uncertainty into a fluent, unqualified assertion.

## 2. Claims are the unit of audit

A claim is a proposition that could be true or false under stated conditions. It is not merely a sentence. One sentence may contain several claims; several sentences may repeat one claim. A title, chart annotation, icon, footnote, or image caption can also make a claim. Editorial review therefore starts by finding falsifiable predicates rather than counting sentences.

Useful signals of a compound claim include:

- conjunctions such as *and*, *while*, *therefore*, or *leading to*;
- causal verbs such as *caused*, *drove*, *improved*, or *generated*;
- comparisons such as *more*, *twice*, *best*, or *faster*;
- quantifiers such as *all*, *most*, *always*, or *up to*;
- population and transport phrases such as *for enterprises* or *across engines*;
- temporal language such as *now*, *within six weeks*, or *long-term*;
- modal language such as *can*, *will*, *is likely to*, or *must*; and
- outcome substitution, where mention becomes citation, citation becomes visibility, visibility becomes traffic, or traffic becomes revenue.

An atomic claim contains one independently testable predicate. This is a working editorial criterion, not a metaphysical definition. “The team deployed schema and citations rose” must become two rows because the deployment could occur without the observed change. “Citations rose from 24 to 41 in the monitored protocol” can remain one row if the numerator, denominator, protocol, and windows are recorded together. “The rise was caused by schema” becomes a separate causal row because it requires a design that addresses alternative explanations.

Atomicity creates two advantages. First, support becomes local. A passage can support the deployment row without being misused for the causal row. Second, revision becomes surgical. If the cross-platform proposition fails, the platform-specific observation may remain. The goal is not to atomize prose forever; it is to establish a governed claim substrate from which accurate prose can later be composed.

## 3. An atomic claim grammar

For W02, normalize a claim into the following fields:

| Field | Audit question |
|---|---|
| Claim ID | Can the proposition be referenced without relying on prose position? |
| Subject | Which entity, document, system, group, or measure is being described? |
| Predicate | What single state, relation, change, comparison, cause, or requirement is asserted? |
| Object/value | What target, category, amount, or state completes the proposition? |
| Population/unit | Over which entities, queries, responses, sessions, or users is it defined? |
| System/surface | Which product, model, interface, API, or benchmark configuration is in scope? |
| Time | What observation, validity, or intervention window applies? |
| Locale | Which language, region, or jurisdiction applies? |
| Conditions | What account state, prompt set, candidate set, or protocol is required? |
| Modality | Is the claim descriptive, comparative, causal, normative, predictive, or prescriptive? |
| Uncertainty | What interval, missingness, disagreement, or unknown state remains? |

The grammar exposes absent scope. “Citations increased” lacks a subject, system, observation window, baseline, denominator, and uncertainty. A stronger descriptive row is: “In the frozen Northstar monitor, on the named en-US answer surface and fixed 50-prompt set, visible source-citation incidence was 24 of 100 recorded opportunities in the January window and 41 of 100 in the February window.” Even that row requires evidence that the parser worked, the runs were comparable, and the prompt set did not change.

Modality deserves special attention. Descriptive evidence does not automatically support a causal verb. A time sequence—change A occurred, then metric B moved—is compatible with causation but also with platform drift, seasonality, query changes, co-interventions, parser changes, sampling variation, and missingness. Converting *was followed by* into *caused* changes the estimand and the required design.

Quantifiers also set a burden. “All engines” requires a defined engine population and observations covering it or a defensible sampling and transport argument. Observing one surface cannot partially support the word *all*; the population claim is out of scope. “Doubled” requires a ratio and a named denominator. Moving from 24 to 41 is a 70.8% relative increase and a 17 percentage-point absolute increase when both denominators are 100. It is neither a doubling nor evidence of an intervention effect by itself.

## 4. Source identity before source use

A source is not just a URL. Before extracting evidence, record its identity:

- issuer, authors, and relevant affiliations;
- source type: paper, standard, official document, repository, tutorial, vendor report, blog, white paper, video, dataset, log, or local artifact;
- title, version, publication status, and date;
- stable identifier and canonical location;
- method, sampling frame, system snapshot, and artifacts where applicable;
- license, access limits, and redistribution status;
- sponsor, commercial interest, or other relevant relationship;
- correction, update, or supersession path; and
- unknown fields, stated as unknown.

Identity is necessary but not sufficient. A DOI verifies an identifier, not a result. An official logo identifies an issuer, not implementation quality. A repository exposes files, not external validity. A national standard can be authoritative about a requirement within jurisdiction and edition while saying nothing about visibility improvement. A polished vendor report can contain a useful observational dataset while lacking a comparator or transparent sampling frame.

Version matters because evidence can decay or change. Platform documentation is bounded by product, surface, locale, and date. A paper preprint may be revised. A benchmark repository may change default parameters. A video may demonstrate an interface state that no longer exists. The claim ledger therefore binds the exact source version to the exact evidence locator and review date. “Current” is not a version.

Independence is also an identity property. Ten pages repeating one press release are not ten independent confirmations. A vendor report summarized by a partner blog and quoted in a webinar remains one evidence family unless another source collected or tested the proposition independently. Record origin, syndication, shared datasets, overlapping authorship, and common funders where material to the claim.

## 5. Evidence ceilings are claim-relative

The course uses seven catalog ceiling codes. They do not form a simple high-to-low ladder. They identify the kinds of claims a source can ordinarily carry within its scope.

`L` covers local controlled artifacts. A file manifest can establish that a package contains six files with specific hashes. It cannot establish that the scientific claims inside are true. `N` covers normative or official authority. A published standard can establish its own requirements and status; it cannot prove that an organization conforms or that a control improves visibility. `A` covers primary research bounded by design, data, models, date, and uncertainty. It can support an experimental or observational result as defined, not a universal production effect.

`B` covers academic pedagogy and maintained research implementations. A respected course or repository can explain a method and provide a reproducible route, while the primary study carries the empirical claim. `C` covers first-party platform documentation for a named product and date. It can establish documented syntax, eligibility, or policy; it does not automatically disclose hidden selection logic or cross-platform behavior. `D` covers vendor observational or quasi-experimental evidence. Such a source may contain valuable operational data, but causal, market-wide, and cross-provider language depends on disclosed design. `E` covers practitioner and secondary synthesis. These sources are useful for terminology, workflow discovery, hypotheses, and contemporary questions, but cannot be the sole authority for scientific, legal, standards, or causal claims.

The ceiling belongs to the **source–claim relationship**, not to the source in the abstract. A first-party specification may be the strongest source for a product’s declared dimensions. The same source is unsuitable as independent evidence that the product outperforms competitors. A peer-reviewed controlled study may support a bounded effect estimate, yet be irrelevant to the current version of an unrelated commercial surface.

Do not average ceilings into a credibility score. A numerical score encourages false compensation: a famous issuer can hide an absent denominator, or many dependent sources can hide one origin. Instead write two sentences: “This source can support…” and “This source cannot establish…”. The second sentence is part of the evidence record, not optional caution text.

## 6. From source to passage to claim

The minimum resolvable path has three layers:

1. **Claim:** the scoped atomic proposition.
2. **Evidence:** the exact passage, table cell, data slice, record, or observation relevant to it.
3. **Source:** the versioned identity that contains or governs the evidence.

A bare source edge—claim to URL—is too coarse. Reviewers need a page, section, paragraph, table, figure, record ID, line range, timestamp for a verified recording, or dataset query. They also need the relation type: supports, partially supports, contradicts, qualifies, is non-informative, or is out of scope.

Entailment is stricter than topical similarity. A document may discuss structured data and citations yet fail to support “FAQ schema caused a 31% sales increase.” Search relevance is not evidential support. The passage must bear on the atomic proposition with compatible objects, population, time, and modality.

Three boundaries should travel with every edge:

- **claim boundary:** what exact proposition is being evaluated;
- **evidence boundary:** what the passage or observation actually shows; and
- **source boundary:** what the source’s method, authority, version, and independence permit.

For example, a deployment record can support “FAQ markup was released to 120 pages on 3 February” if it names the commit, page set, and date. Its evidence boundary is the recorded deployment. Its source boundary is an internal operational record. It is non-informative about citation effect unless the record also contains a suitable outcome design—which deployment records ordinarily do not.

## 7. Dispositions, contradiction, and unresolved states

W02 uses five primary dispositions:

- **Supported:** the available evidence entails the claim within compatible scope, and no unresolved critical conflict is known.
- **Partially supported:** a narrower component is supported, but at least one scope, quantifier, modality, or value must be revised.
- **Contradicted:** relevant evidence supports an incompatible proposition under comparable scope.
- **Unresolved:** relevant evidence conflicts, is too incomplete, or cannot distinguish important explanations.
- **Out of scope:** the evidence does not observe or govern the population, system, time, jurisdiction, or outcome named by the claim.

Unsupported can be used as a release label, but it is helpful to explain why: no evidence edge, locator missing, stale version, source dependence, construct mismatch, design mismatch, or unresolved contradiction. “No evidence found” is not the same as “false.” It records the current search and review boundary.

Contradictions must not be deleted to make the ledger clean. First compare claim scope. Two numbers may use different denominators, windows, locales, or definitions. Next compare source versions and methods. One may be newer, one may sample a different population, or both may be biased in different ways. Then decide whether the claim can be narrowed, whether one edge is invalid, or whether the state remains unresolved. Record the adjudication, reviewer, date, and excluded alternative.

Preserving disagreement supports later correction. Silent selection creates a brittle knowledge base: the discarded result returns without its history, and future editors cannot see why it was rejected. An unresolved row with a next evidence action is often stronger than a confident row built by omission.

## 8. Five over-generalization tests

Before expression, apply five transport tests.

**Population transport.** Does a claim move from sampled prompts, documents, categories, users, or organizations to a broader population? Name both. A convenience sample of controversial questions is not all information needs.

**System transport.** Does evidence from one model, API, interface, benchmark, or retrieval configuration become a claim about other systems? “Across AI engines” is especially demanding because engine, surface, and version must be defined.

**Time transport.** Does a historical snapshot become a present-tense property? Model and platform behavior can drift; documentation and policies change. State the observation and review dates.

**Causal transport.** Does an association or before/after difference become an intervention effect? Require a comparator and a design appropriate to confounding, interference, repeated measurements, and drift.

**Outcome transport.** Does one event become another? Mention is not citation. Citation is not entailment. Entailment is not absorption. Visibility is not referral. Referral is not conversion. Assisted conversion is not revenue caused by the content change.

Each test asks the same question: what bridge would be needed to move from the measured object to the expressed claim? If the bridge is absent, narrow the claim or define the next study. Do not fill the gap with tone.

## 9. Quantitative sentences need named nouns

Percentages are incomplete without a noun and denominator. “Citations increased 71%” might mean citation count, share of responses containing at least one citation, citations per answer, distinct cited domains, or monitored opportunities. Those are different metrics. Report the numerator, denominator, unit, aggregation, missingness, and uncertainty.

Relative and absolute change answer different questions. From 24 of 100 opportunities to 41 of 100 is a 17 percentage-point increase and a 70.8% relative increase. If the monitoring windows have different run counts, raw counts are not directly comparable. If four prompts changed, the estimand is no longer the fixed-prompt change unless the analysis removes or separately reports them.

A percentage can also hide multiplicity. If a team tested many prompts, platforms, time windows, and content variants and reported only the largest gain, the selected result does not have the uncertainty of a preregistered single comparison. W02 does not require a full statistical model for every operational claim, but it requires the ledger to show which comparisons were planned, which were exploratory, and what was omitted.

Commercial outcomes require construct discipline. Four assisted transactions after a change may be a 30.8% increase from 13 to 17, but the attribution rule may count any session touching an AI referral within 30 days. That number is not “AI-generated sales,” and a before/after increase is not a causal effect. The safe sentence names the operational definition and avoids the causal verb.

## 10. Reading P20 and P40 without title transfer

P20, *What Evidence Do Language Models Find Convincing?*, studies a controlled conflict task in which language models answer controversial binary questions from two opposing passages. In the reviewed version, model-specific filtering creates different effective samples; source identity cues are removed; and relevance-oriented counterfactual edits are more consistently associated with stance adoption than the evaluated style changes. This route is useful for asking what passage properties can influence a bounded model decision.

Its evidence ceiling does not extend to factual truth, source credibility, open-web retrieval, production citation, present-day multi-source synthesis, or human persuasion. A passage “winning” means the model’s answer adopts its stance in that setup. It does not mean the passage is true or ethically preferable. The study therefore teaches a boundary: model sensitivity to semantic relevance must not be rewritten as a universal content prescription or authority signal.

P40, *GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning*, compares verification procedures while replacing zero, one, two, or three items in a fixed three-evidence set with synthetic or tampered material. It is useful for showing that degradation is attack-specific and that clean performance does not settle poisoned robustness. It also motivates preserving source families, conflict, and residual uncertainty in a claim graph.

The reviewed paper reports one run per fixed configuration, incomplete release information, limited annotation detail, and a controlled benchmark that does not estimate real poisoning prevalence. Its point estimates cannot establish a stable universal method ranking or production fact-checking reliability. W02 uses P40 as a threat to naïve aggregation: more or apparently corroborating evidence can be engineered, dependent, or misleading. It does not use the paper as proof that a real source set has been poisoned.

## 11. The evidence-first workflow

An operational review can proceed in twelve steps:

1. Freeze the original expression and preserve its location.
2. Underline quantifiers, causal verbs, comparisons, systems, populations, times, and outcomes.
3. Split the expression into atomic propositions.
4. Classify each proposition’s modality.
5. Create a source identity record for every candidate source.
6. Cluster dependent sources by origin and shared data.
7. Extract exact evidence locators without rewriting them into stronger language.
8. Assign a claim-relative ceiling and write “can support / cannot establish.”
9. Type each evidence edge and preserve qualification or contradiction.
10. Apply the five transport tests and quantitative-noun check.
11. Choose a release action: approve, narrow, caveat, seek evidence, escalate, or block.
12. Record reviewer, date, unresolved issue, and next evidence action.

This workflow is intentionally reversible. A future source can be added without losing the original claim or earlier judgment. A changed platform document can trigger review of affected rows. A contradiction can move a claim from approved to unresolved without erasing the previously used evidence.

The release gate is simple: a public claim must have a resolvable evidence path within its ceiling, or it must carry an explicit unsupported status and be withheld or clearly labeled. “There is a citation somewhere in the document” is not a release condition.

## 12. Worked synthetic case preview

The Northstar case supplies nine fictional source cards: a release note, a deployment manifest, a monitoring export, a parser quality note, a change log, a CRM report, a platform help card, a practitioner tutorial, and an internal methods memo. Their variety is deliberate. Some are strong for deployment facts, some for a bounded observation, some only for terminology or hypothesis generation, and none alone supports the headline’s causal and cross-platform language.

Learners atomize six claims and map evidence. The correct resolution preserves useful facts: a documented deployment occurred; a particular monitored citation incidence changed; and AI-assisted transactions changed under a stated attribution rule. It rejects or suspends the stronger claims: the change did not double, only one surface was measured, multiple co-interventions occurred, prompt and parser differences remain, and the CRM metric does not identify sales caused by schema.

The publication-safe rewrite is less dramatic and more informative. It names the synthetic protocol, both denominators, the absolute change, the co-interventions, and the inability to attribute causality. This is not timid writing. It is writing that lets a second analyst reconstruct the statement.

## 13. Discussion prompts

1. Which source in your field is highly credible yet unsuitable for a claim people commonly attach to it?
2. When should “partially supported” be replaced by two rows, one supported and one unsupported?
3. How would you distinguish a genuine contradiction from different denominators or time windows?
4. What evidence would be sufficient to move the Northstar causal claim from unresolved to supported in its experimental environment?
5. When does source dependence matter more than source count?
6. Which fields should trigger automatic re-review when a platform or standard changes?
7. What is lost when a writer removes uncertainty from a headline but restores it only in a footnote?

## 14. Summary and exit test

Evidence before expression is a data model for accountable writing. It separates the atomic claim from the evidence item and the source identity. It treats authority as claim-relative, not as a prestige score. It binds every support decision to scope, version, exact location, and source family. It preserves contradiction and unresolved states. It tests population, system, time, causal, and outcome transport before a sentence is released.

For the exit test, take one compound statement and produce:

1. at least four atomic claims;
2. one complete source identity record;
3. one exact evidence edge;
4. one credible but unsuitable source with the reason;
5. one disposition of supported, partially supported, contradicted, unresolved, or out of scope; and
6. one publication decision plus the cheapest ethical next evidence action.

If a second reviewer cannot reconstruct why the claim was approved or blocked, the audit is not finished. If the evidence cannot support the expression, revise the expression—not the evidence boundary.
