# W02 Worked Case — The Northstar Atlas Headline

## Status, safety, and use

**SYNTHETIC TEACHING FIXTURE.** Northstar Atlas, Aster Answer, all people, all documents, and all observations in this case are fictional. The fixture was written for offline teaching. It contains no measurement of a real company, platform, product, crawler, customer, or user. Numerical values are synthetic and were selected to expose claim, denominator, and design errors—not to approximate an industry effect.

Use the case in two passes. Learners first receive Sections 1–5 and build their own ledger. The model resolution begins in Section 6. An instructor should not present the resolution as a hidden ground truth about how real generative systems work; it is an adjudication of the supplied synthetic record.

## 1. Publication candidate

The fictional Northstar Atlas marketing team proposes this headline:

> After Northstar Atlas added FAQ schema, citations across all AI engines doubled and AI-generated sales rose 31% in six weeks.

The proposed supporting paragraph says:

> The February optimization made Northstar the preferred source for generative answers. Monitoring proves that structured data drove citation growth, while the customer dashboard confirms a direct revenue effect. The same evidence shows that the method works across providers.

**Learner task:** preserve the original wording, underline every quantifier, causal verb, comparison, system, time window, and outcome, then split the headline into exactly six atomic claim candidates. Do not improve the prose yet.

## 2. Frozen source bundle

Each card gives an identity, exact evidence excerpt, ordinary course ceiling, and source boundary. A ceiling constrains use; it does not decide the claim disposition automatically.

### SC-01 — Change-request release note

- **Identity:** Northstar internal change request `NS-218`, revision 4, approved 2 February 2026, exported as a frozen PDF.
- **Family:** F1 · internal operations.
- **Ceiling:** `L` — establishes the state of a controlled local record.
- **Locator `SC-01 §2, “Release scope”`:** “Target deployment: 120 English help-center pages. Bundle: FAQ JSON-LD, revised page headings, updated answer summaries, and related-article links. Planned release: 3 February 2026.”
- **Can support:** the approved target, planned date, and listed bundle.
- **Cannot establish:** successful deployment to every page, external eligibility, citation change, causal effect, cross-platform behavior, or commercial outcome.

### SC-02 — Deployment manifest

- **Identity:** signed local manifest `deploy-2026-02-03.json`, build `7f21c9`, generated 3 February 2026 at 18:10 UTC.
- **Family:** F1 · internal operations.
- **Ceiling:** `L`.
- **Locator `SC-02 records.summary`:** `targeted=120; completed=118; rolled_back=2; faq_jsonld_valid=116; heading_revision=72; answer_summary_revision=40; related_links_revision=88`.
- **Locator `SC-02 records[44,97]`:** two pages were rolled back after validation errors and remained on the January version.
- **Can support:** recorded deployment and validation states for the frozen build.
- **Cannot establish:** that public crawlers fetched the pages, that a platform accepted the markup, or that any observed outcome was caused by the bundle.

### SC-03 — Citation-monitor export

- **Identity:** Northstar monitoring export `monitor-v2.3.csv`, generated 2 March 2026; one fictional answer surface named Aster Answer; locale `en-US`.
- **Family:** F2 · monitoring instrumentation.
- **Ceiling:** `L/D` — a controlled export of an internal vendor-style observational monitor.
- **Locator `SC-03 README, rows 8–16`:** January window: 50 prompt IDs, two runs per prompt, 100 response opportunities, 24 responses with at least one visible citation to a Northstar help page. February window: 50 prompt IDs, two runs per prompt, 100 response opportunities, 41 such responses.
- **Locator `SC-03 README, rows 18–20`:** prompts `Q07`, `Q19`, `Q33`, and `Q48` received wording corrections before the February run; no paired analysis excluding them is included.
- **Can support:** the values present in the export and its declared protocol.
- **Cannot establish:** perfectly measured citations, a fixed-prompt comparison, an effect of FAQ schema, other surfaces, user behavior, or revenue.

### SC-04 — Parser quality note

- **Identity:** monitoring QA note `parser-check-2026-02.md`, reviewed 4 March 2026; stratified sample of 40 rendered responses from both windows.
- **Family:** F2 · monitoring instrumentation.
- **Ceiling:** `L`.
- **Locator `SC-04 Table 1`:** against manual review, the parser produced 2 false-positive Northstar citation labels and 1 false-negative label in the 40-response sample. Errors were not propagated back to the full export.
- **Locator `SC-04 Limitations`:** interface layout changed on 15 February; no separate validation is reported before and after that change.
- **Can support:** observed QA errors in the inspected sample.
- **Cannot establish:** the corrected full-window citation incidence or the direction and magnitude of total parser bias.

### SC-05 — Site and communications change log

- **Identity:** internal weekly operations log, frozen 28 February 2026.
- **Family:** F1 · internal operations.
- **Ceiling:** `L`.
- **Locator `SC-05 Week 6`:** FAQ markup bundle released; 72 headings and 40 answer summaries revised; 88 related-article link blocks changed.
- **Locator `SC-05 Week 7`:** median server response time reduced from 680 ms to 410 ms after cache configuration; two trade publications linked to the help center following a product announcement.
- **Locator `SC-05 Week 8`:** Aster Answer interface revision observed; monitoring parser hotfix proposed but not shipped during the window.
- **Can support:** recorded co-interventions and concurrent events.
- **Cannot establish:** the effect of any one change or whether the events changed retrieval, generation, citation, or sales.

### SC-06 — CRM attribution report

- **Identity:** Northstar internal CRM dashboard export, generated 5 March 2026; attribution rule version `assist-30d-v1`.
- **Family:** F3 · commercial analytics.
- **Ceiling:** `L/D`.
- **Locator `SC-06 Table A`:** January: 260 sessions with an `ai_referral` tag and 13 transactions with any such session in the prior 30 days. February: 390 tagged sessions and 17 assisted transactions.
- **Locator `SC-06 Methods`:** the tag is assigned by a referrer-domain list; direct visits after an earlier tagged session remain assisted; transaction revenue and margin are not included; no page-level citation, prompt, or answer-surface record is joined to the CRM table.
- **Can support:** the dashboard counts under its stated attribution rule.
- **Cannot establish:** “AI-generated sales,” incremental sales, revenue, profit, a link to visible citations, or an effect of the February site changes.

### SC-07 — Aster Answer publisher help card

- **Identity:** fictional first-party help card, version 2026-02-10, “Structured data and answer features.”
- **Family:** F4 · platform documentation.
- **Ceiling:** `C` — first-party documentation for the named surface and version.
- **Locator `SC-07 ¶3`:** “Supported structured data can help systems interpret eligible page elements. Eligibility does not ensure retrieval, inclusion in a generated answer, or visible source attribution.”
- **Locator `SC-07 ¶6`:** “Feature behavior may vary by query, locale, interface, and system update.”
- **Can support:** what Aster Answer publicly declares about the named feature in this fictional version.
- **Cannot establish:** hidden ranking weights, successful processing of Northstar pages, treatment effect, another provider’s behavior, or commercial outcome.

### SC-08 — Practitioner tutorial

- **Identity:** fictional agency blog, “FAQ schema for AI citations,” published 12 January 2026; author and three named client anecdotes; no downloadable data.
- **Family:** F5 · practitioner synthesis.
- **Ceiling:** `E`.
- **Locator `SC-08 heading “Results”`:** “In our recent launches, adding FAQ schema was followed by more citations and more qualified leads.”
- **Locator `SC-08 footer`:** the post supplies no query set, run count, parser rule, comparison window, control, missingness, uncertainty, or client-level results.
- **Can support:** a practitioner’s stated workflow and hypothesis.
- **Cannot establish:** the occurrence or cause of Northstar’s changes, a scientific effect, a cross-platform effect, or a sales effect.

### SC-09 — Measurement methods memo

- **Identity:** Northstar internal memo `measurement-boundary-v1.md`, approved 6 March 2026.
- **Family:** F1 · internal methods and governance.
- **Ceiling:** `L`.
- **Locator `SC-09 §1`:** “The monitor covers Aster Answer only. No other engine or surface was sampled.”
- **Locator `SC-09 §2`:** “The January–February comparison was exploratory, not preregistered. There was no concurrent control page set, random assignment, interrupted-series design, or drift block.”
- **Locator `SC-09 §3`:** “FAQ markup, headings, answer summaries, links, latency, publicity, four prompt wordings, and the monitored interface changed within the study period.”
- **Locator `SC-09 §4`:** “The CRM and monitoring datasets have no common response, session, page, or user key.”
- **Can support:** the documented limits of the local analysis.
- **Cannot establish:** the magnitude of bias, the counterfactual outcome, or which co-intervention matters.

## 3. Learner worksheets

### Worksheet A — Source identity and dependence

For each source, record issuer, type, version/date, family, method, artifact, ceiling, exact locator, license/redistribution status, correction path, and unknowns. Then answer:

1. Why are SC-01, SC-02, SC-05, and SC-09 not four independent confirmations?
2. Why is SC-07 authoritative for one declared product behavior but non-informative about Northstar’s causal effect?
3. What value does SC-08 provide even though it cannot carry the headline?
4. Which source controls the meaning of “assisted transaction,” and which source does not?

### Worksheet B — Atomic claim ledger

Create six rows with these fields:

`claim_id`, `original_span`, `atomic_claim`, `modality`, `subject`, `predicate`, `object_or_value`, `population_or_unit`, `system_surface`, `time`, `locale`, `conditions`, `source_ids`, `exact_locators`, `ceiling`, `edge_type`, `disposition`, `conflict`, `release_action`, `next_evidence_action`.

### Worksheet C — Quantitative check

Calculate and name the noun for each result:

1. absolute percentage-point change in exported visible-citation incidence;
2. relative change in that exported incidence;
3. relative change in assisted-transaction count;
4. session-to-assisted-transaction ratio in each month; and
5. whether any supplied value is transaction revenue or incremental sales.

### Worksheet D — Release decision

Choose one action for every row: approve as written, approve after narrowing, publish with a visible caveat, seek specified evidence, escalate for specialist review, or block. A block is a valid completed decision.

## 4. Required claim/evidence/source graph

Create three node bands:

- top band: six `CLM` nodes;
- middle band: exact evidence locators, not whole documents; and
- bottom band: nine `SC` source identities grouped into five families.

Use a solid line labeled **supports**, a dotted line labeled **partial**, a double line labeled **contradicts**, and a broken grey line labeled **out of scope**. Shape and text must duplicate color. Every evidence node includes the measured unit or governing scope. The graph is a teaching model; it is not a representation of a real platform architecture.

## 5. Stop point before the model resolution

Do not continue until the learner has:

- exactly six atomic rows;
- at least one exact locator on every row that has evidence;
- at least one contradicted or unresolved row;
- an explicit source-family judgment;
- distinct labels for descriptive and causal claims; and
- a release action for all six rows.

## 6. Model atomization and adjudication

| ID | Atomic claim | Best evidence edge | Ceiling | Disposition | Release action |
|---|---|---|---:|---|---|
| `CLM-01` | FAQ schema was successfully deployed to all 120 targeted help pages on 3 February 2026. | SC-01 §2 gives the target; SC-02 summary records 118 completed, 2 rolled back, and 116 valid. | `L` | **Contradicted as written; partially supported after narrowing.** | Replace with the recorded counts and distinguish target, completed, and valid states. |
| `CLM-02` | Visible citations doubled between the January and February monitored windows. | SC-03 README reports 24/100 and 41/100; SC-04 records parser errors. | `L/D` | **Contradicted.** | Do not use “doubled.” If needed, report the export values with QA and prompt-change caveats. |
| `CLM-03` | FAQ schema caused the observed increase in visible citation incidence. | SC-05 and SC-09 document co-interventions, no control, prompt changes, drift, and interface change. | `L` for design record | **Unresolved and unsupported for release.** | Block the causal verb; design a controlled, drift-aware study. |
| `CLM-04` | The citation increase occurred across all AI engines. | SC-09 §1 states that only Aster Answer was monitored. | `L` | **Out of scope and contradicted by the measurement scope.** | Remove the cross-engine claim. A broader study needs a defined platform sample and transport rule. |
| `CLM-05` | Under `assist-30d-v1`, assisted-transaction count increased from 13 in January to 17 in February, approximately 31%. | SC-06 Table A and Methods. | `L/D` | **Supported as a dashboard description within its operational definition.** | Publish only if the operational definition, raw counts, windows, and non-causal status are visible. |
| `CLM-06` | The FAQ-schema change caused 31% more sales within six weeks. | SC-06 has assisted counts but no revenue; SC-09 says datasets are unlinked and design is non-causal. | `L/D` | **Unsupported; construct and causal mismatch.** | Block. Require a defined sales metric, joinable exposure, comparator, and causal design. |

The six rows show why one citation beside the headline is inadequate. SC-02 is strong for a deployment record and irrelevant to commercial causality. SC-06 is strong for its own dashboard counts and irrelevant to a source-citation mechanism. SC-07 is the official source for Aster’s declared feature boundary and cannot establish that Northstar received a benefit.

## 7. Model calculations

### Citation incidence

- January exported incidence: `24 / 100 = 0.24`, or **24%**.
- February exported incidence: `41 / 100 = 0.41`, or **41%**.
- Absolute change: `41% − 24% = 17 percentage points`.
- Relative change: `(41 − 24) / 24 = 0.7083`, or approximately **70.8%**.
- Doubling test: doubling 24 would require 48 under the same denominator; 41 is not 48.

These calculations describe uncorrected export labels. They do not repair the parser, restore four changed prompt wordings, isolate the interface change, or estimate a causal effect.

### Assisted transactions

- Count change: `17 − 13 = 4` transactions.
- Relative count change: `4 / 13 = 0.3077`, or approximately **30.8%**.
- January assisted transactions per tagged session: `13 / 260 = 5.0%`.
- February assisted transactions per tagged session: `17 / 390 ≈ 4.36%`.

The ratio is not automatically a conversion rate because the unit and attribution rule may allow a transaction and multiple sessions to interact in ways not reported. More importantly, none of these values is revenue, profit, incremental sales, or a causal effect.

## 8. Invalid evidence moves

1. **Citation laundering:** attach SC-01 to the whole headline because it proves that a change was planned.
2. **Prestige substitution:** use SC-07’s first-party status as proof of Northstar’s outcome.
3. **Topical entailment:** use SC-08 because it contains the words *schema*, *citations*, and *leads*.
4. **Source-count inflation:** count four F1 documents as four independent confirmations.
5. **Temporal causality:** infer that the deployment caused every later change.
6. **Platform transport:** replace one named surface with “all AI engines.”
7. **Outcome substitution:** rename assisted transactions as AI-generated sales or revenue.
8. **Relative-only framing:** report 70.8% without the 24/100 and 41/100 counts.
9. **Error suppression:** omit SC-04 because its parser caveat weakens the narrative.
10. **Conflict deletion:** keep the 120-page target and discard the 118/116 deployment states.

## 9. Bounded rewrite

One publication-safe internal summary is:

> In this synthetic exploratory monitor, Aster Answer responses with an exported visible citation to a Northstar help page increased from 24 of 100 recorded opportunities in January to 41 of 100 in February, an uncorrected change of 17 percentage points. Four prompt wordings, the monitored interface, the parser context, site content, internal links, latency, and publicity also changed. The record therefore does not identify an effect of FAQ schema or support transport to other answer surfaces. Separately, the fictional CRM counted 13 and 17 transactions with an AI-referral-tagged session in the prior 30 days; that operational metric is not incremental sales and is not joined to the citation monitor.

This rewrite is longer than the headline because it restores the objects that the headline collapsed. A shorter external statement could say that an exploratory monitor showed a change that requires controlled follow-up, but it should not name a schema effect, cross-engine result, or sales effect.

## 10. Cheapest ethical next evidence actions

| Gap | Next action | What the action would and would not add |
|---|---|---|
| Deployment state | Repair or document the two rollbacks; validate all 120 pages; preserve hashes | Establishes technical state, not platform processing or effect |
| Parser error | Double-code a larger stratified sample across interface versions; report disagreement and corrected sensitivity bounds | Improves measurement, not causal identification |
| Prompt comparability | Re-run the fixed original prompt set and separately report the corrected set | Restores one comparison boundary, not platform stability |
| Schema effect | Use a preregistered page-level or cluster-level intervention with a suitable control, blocked repetitions, and locked co-interventions | Estimates a local effect under that design, not every engine or future version |
| Cross-platform transport | Define the platform/surface population, sample named configurations, and report heterogeneity without pooling away failures | Broadens observed scope, not an unrestricted universal claim |
| Commercial effect | Define transaction, revenue, and attribution; lawfully link exposure units; use a design appropriate to confounding and interference | Can address an incremental outcome under conditions, not automatic long-term value |

## 11. Self-check and handoff

A learner’s case is ready for handoff when another reviewer can answer:

- Which words in the original expression became separate claims?
- Which exact evidence locator supports or contests each row?
- What is the source family, ceiling, and version?
- Which edges are non-informative despite topical relevance?
- Which contradictions and unknowns remain?
- Why is each release action proportionate to the evidence?
- What new evidence would change the disposition?

The final handoff consists of the six-row ledger, source identity table, graph or printable matrix, calculation note, conflict log, bounded rewrite, and next-action table. The successful result is not the most positive claim. It is the strongest expression that the frozen evidence can actually carry.
