1. The opening question
Suppose an answer engine names a product, cites its documentation, and places it first in a recommendation. What became visible? The entity name appeared. A source link appeared. Several claims appeared near that link. The product occupied an early span in the answer. Perhaps a user then clicked, asked a follow-up question, or acted. These are related events, but they are not the same event. They do not have the same unit, denominator, causal meaning, or evidence requirement.
This distinction is the starting point for generative engine optimization (GEO). A weak formulation begins with a desired tactic—rewrite a heading, add a quotation, publish more pages—and then searches for a favorable number. A rigorous formulation begins by defining the research object. It asks which source, which query population, which system surface, which stage, which visible event, which time window, and which comparison are under study. Only then can an intervention be described and evaluated.
The object-first discipline matters because the visible answer is an interface, not a transparent record of the machinery behind it. A system might rewrite a query, search one or more corpora, retrieve candidates, rerank them, allocate a finite context, generate language, and attach citations. It might instead answer partly from parametric memory. It might repeat some operations or omit others. A commercial surface may reveal only the final response. The scientific task is therefore not to fill hidden stages with a persuasive story. It is to state what was observed, what was inferred, what remains a hypothesis, and what evidence could distinguish the alternatives.
2. Defining the GEO research object
For this course, GEO is the study and responsible intervention of source visibility, use, attribution, and downstream effects in generative search and recommendation systems. The definition contains four constraints.
First, GEO concerns governed information sources, not merely strings. A source has an issuer, version, date, method, claims, permissions, and correction path. Second, visibility is plural: mention, citation, source use, factual fidelity, prominence, and action can move in different directions. Third, the system and surface must be named. A web answer, file-grounded chat, enterprise RAG application, and recommendation agent expose different variables. Fourth, an intervention is responsible only when it preserves evidence, factual meaning, authorization, accessibility, and declared risk limits.
A bounded study can be represented by a scope tuple
where:
- is the source or source population;
- is a declared query distribution rather than an opportunistic prompt list;
- is the named system or model state that can actually be recorded;
- is the user-facing surface and configuration;
- is the observation window;
- is the target event and its unit;
- is the comparator or counterfactual design; and
- is the evidence and safety boundary.
“Improve AI visibility” leaves every slot ambiguous. A bounded alternative is: “For the supplied 40-query campus-mobility panel, on the frozen open retrieval sandbox, does an evidence-equivalent heading treatment change the proportion of runs in which the HarborCell H2 passage enters the top-five candidate set, relative to the unchanged page, during version 1.2 of the fixture, without changing any product claim?” This question still requires a protocol, but it names an event that can be falsified.
GEO overlaps with adjacent practices without collapsing into them. Search engine optimization often observes indexed pages, ranked links, and visits. Answer-engine optimization often focuses on extraction into direct answers. RAG engineering controls a corpus, chunking, retrieval, context, and generation. Digital public relations works through legitimate third-party publication and relationships. GEO can draw methods from each, but a high-ranked link need not enter a generated context; a citation need not indicate faithful source use; and repeated coverage need not be independent corroboration. The boundaries should be drawn by research objects, not job titles.
The reading route reinforces this framing without serving as a shortcut. P16 helped establish visibility within a generated response as an explicit optimization object, but its principal treatment modifies a source already inside a fixed top-five context. That evidence cannot be used to infer open-web discovery, ordinary retrieval, traffic, or revenue. P19 is a position paper that organizes risks and possible intervention surfaces; it does not experimentally establish prevalence, harm, or provider effectiveness. P06 is a narrative map of search–LLM integration, not a systematic review or platform blueprint. P15 describes heterogeneous August 2025 API snapshots; it does not test a content intervention or establish current behavior. These boundaries are part of the lesson, not footnotes to remove during presentation.
3. The objects that must remain separate
A GEO record commonly contains at least nine object types:
- Entity: an organization, product, person, location, or concept whose identity must be resolved.
- Claim: a minimal proposition with scope, time, locale, and conditions.
- Source: a governed document, dataset, record, page, or other information object that can support or contest claims.
- Representation: a passage, chunk, image description, embedding, index entry, or transformed unit derived from a source.
- Query: a measurement instrument drawn from a defined population or intentionally constructed test set.
- Response: the generated language and visible organization presented on a named surface.
- Attribution: a displayed relation such as a citation marker, link, footnote, source panel, or provenance record.
- Event: a predeclared occurrence, such as an entity mention, resolved source citation, supported proposition, or click.
- Outcome: the consequence valued by a stakeholder, which may be an event or an aggregate over events.
Why insist on this vocabulary? Consider the statement “The brand gained citations.” The subject may be an entity, while a citation normally points to a source. A source-domain citation count can increase even if the brand name never appears. Conversely, the brand can appear with no citation. If a report alternates between entity and source as the unit, its numerator loses meaning.
Claims also require atomization. “The HarborCell H2 is a safe, long-range battery with a 36-month warranty” contains at least three propositions: a safety proposition, a range proposition, and a warranty-duration proposition. Each may need a different source and scope. One citation at the end of the sentence cannot automatically support all three. The evidence graph must connect each atomic claim to the exact passage and source identity that bears on it.
The same rule applies to outcomes. A favorable mention can be inaccurate. A faithful citation can be visually obscure. A prominent recommendation can generate no action. A visit can occur for a reason unrelated to the generated answer. Aggregating these into one score hides trade-offs and creates opportunities for selective reporting. W01 therefore treats the vector of events as primary; composite objectives belong later, after policy weights and non-compensable constraints are declared.
4. The six-stage conditional visibility chain
The course uses a six-stage diagnostic chain:
- Discoverable: the source is available to the relevant acquisition process.
- Retrieved: the source or representation appears in a candidate set for the query.
- In context: the source survives selection and enters the effective generation context.
- Contributes: source information shapes generated content.
- Attributed: the contribution receives a visible and sufficiently faithful source relation.
- Outcome: the declared user, source, platform, or social consequence occurs.
Let the corresponding events be . For an outcome that requires every event, the joint probability can be decomposed by the chain rule:
This expression introduces no independence assumption. It makes dependence visible. A revision could help a passage match a query while making the answer less faithful. A longer evidence section could improve support but exceed a context budget. A source could contribute to an uncited sentence. An interface could display a citation chosen after generation. The direction of one transition does not determine the others.
The chain is a research abstraction, not a disclosure that every platform implements six modules. A system may combine retrieval and context selection, conduct several searches, use tools after an initial response, or answer without a fresh retrieval event. In a file-grounded sandbox, candidate retrieval and context may be logged directly. On a closed surface, they may remain hidden. Learners should alter the diagram when a studied system differs; forcing observations into the diagram would defeat its purpose.
Every arrow needs a denominator. “Retrieved in 18 cases” is incomplete without the eligible query count, failure policy, repeated-run rule, and definition of retrieval. “Cited 45 percent of the time” is incomplete if the denominator switches between all queries, successful responses, responses with any citation, and source-eligible responses. A meaningful stage map writes the denominator under the arrow and marks it “unknown” when the system does not expose it.
5. Observable and latent components
Observability is protocol-relative. A variable is not simply observable or hidden in the abstract; it is observable under a stated collection design.
- Surface-observable: visible response text, displayed citation target, ordering, timestamp recorded by the collector, and permitted interaction trace.
- Controlled-system observable: candidate scores, rank positions, selected passages, prompt templates, model checkpoints, and random seeds in an open or frozen sandbox.
- Proxy-sensitive: a construct such as source contribution inferred from distinctive claims, perturbation, or attribution, where alternative explanations remain.
- Latent in the study: hidden candidate sets, proprietary ranking features, internal model activations, training inclusion, or unrecorded personalization state.
A screenshot is strong evidence for a narrow visible-state proposition: at the recorded time, the captured surface displayed the represented text and citation markers. It is weak evidence for frequency because it supplies one realization and often no sampling frame. It is not direct evidence of the candidate pool, ranking function, training corpus, or causal reason for the response.
Good records use an epistemic-status vocabulary. An observation reports a recorded event. An interpretation assigns meaning using stated definitions. A mechanism hypothesis proposes a process that would explain observations and names a discriminating test. A causal effect claim requires a comparison design and assumptions. An unsupported promise asserts an outcome beyond the available evidence. These labels are not grades of enthusiasm; they are commitments about what could make the statement wrong.
The open–closed distinction also limits transfer. An open sandbox is valuable because it exposes intermediate traces and supports ablation. A closed surface is valuable because it shows externally visible behavior in a real interface. The former does not automatically represent the latter. The latter does not reveal the former’s mechanisms. A defensible project often uses both: open systems to test mechanisms and bounded field observations to study surface behavior, with no silent bridge between them.
For each observation, assemble a minimal evidence packet: system and surface name, caller-visible configuration, date and timezone, locale and language, account/session state when allowed, exact query, collection order, visible response, resolved citation targets, failure state, collection method, and checksum of the frozen artifact. Unknown system properties remain fields marked unknown. Filling them with a familiar architecture diagram creates false precision.
6. Measurement objects and denominators
W01 separates six measurement families.
Mention uses an entity–response unit. A positive event means a resolved entity or alias occurs in a response. It does not establish source use, accuracy, recommendation, or favorable sentiment. The denominator might be all valid responses to a query panel, but entity-resolution and missing-response rules must be stated.
Citation uses a source–response unit. A positive event means a displayed attribution resolves to the governed source or a declared source family. It does not establish that the adjacent claim is entailed, that the source caused the wording, or that a user noticed the link. Redirects, duplicate URLs, domain aggregation, and citation-panel behavior require explicit handling.
Absorption uses a claim–source–response unit. It asks whether source-distinctive evidence is reflected in the answer. Common facts cannot identify a particular source. A strong absorption protocol segments claims, records exact supporting passages, searches plausible alternative sources, and retains an unresolved category. Visible citation and absorption should be scored separately.
Fidelity uses a claim–response unit. It measures whether the generated proposition preserves the scoped meaning of the reference claim. A response can cite the correct page while changing a number, omitting a condition, or turning an association into a causal statement. Fidelity evaluation therefore operates at claim level, not citation-count level.
Prominence uses a span–response unit. It may encode order, normalized position, share of answer space, or exposure under a declared reading model. Early position is not automatically attention, favorable treatment, or utility. The function must be specified before observing results.
Action uses a user–task or event-log unit. A click, follow-up, purchase, or tool action is downstream of the answer, but attribution requires consent, logging, and an identification design. An action record alone does not reveal why the user acted.
A metric card should contain: construct, event rule, unit, eligible population, numerator, denominator, query weighting, repeated-run rule, missingness categories, aggregation, uncertainty method, and prohibited interpretation. If two dashboards use “citation share” with different denominators, they define different quantities even when the labels match.
7. Evidence ceilings
An evidence ceiling states the strongest claim class an item can ordinarily support after identity and method checks. It is claim-relative, not a prestige ladder.
- E0 — public assertion: promotional pages, uncited posts, or clips can establish what an actor publicly asserted. They do not validate the assertion.
- E1 — declared rule or status: official documentation, policy, and standards can support named behavior, syntax, status, or requirements within scope. They do not expose hidden weights or establish intervention efficacy.
- E2 — inspectable procedure: method-rich tutorials, white papers, and code walkthroughs can support how a disclosed procedure is implemented. Their examples do not establish general effects.
- E3 — bounded observation or benchmark: a documented observational study or benchmark can support associations and performance under its sample, data, systems, date, and metric.
- E4 — controlled comparison: an intervention with a defensible comparator can support an effect inside the experimental environment, subject to design assumptions.
- E5 — stronger transport evidence: independent replication across settings or a strong field design can support a more transportable estimate, while heterogeneity and scope still remain.
Identity precedes ceiling. Record issuer, item type, title, version, date, status, method, artifacts, license, and stable route. A repository can expose code but not independently validate a paper. A standard can state a control but not prove an organization implements it. A first-party platform page can define the named feature it documents but not cross-engine behavior. A paper can support only the claims its design identifies; a confident title is not an estimand.
Ceilings also apply to visuals. An attractive pipeline in a paper may be a conceptual drawing rather than an observed trace. Before reuse, check version, authorship, figure number, license, modification status, and accessible alternative. When permission is uncertain, cite the concept and build an original diagram whose caption states the synthesis and boundary.
8. Worked synthetic case: HarborCell H2
The complete data and audit appear in WORKED_CASE.md; this section demonstrates the reasoning.
The fictional manufacturer Blue Marsh Mobility publishes a versioned page for the fictional HarborCell H2 battery. The page states a nominal energy capacity of 480 Wh, an enclosure rating of IP54 under specified conditions, and a 24-month limited warranty. It does not present comparative safety testing, estimated range, or a 36-month warranty. A frozen synthetic answer surface responds to the query “Which campus e-bike battery is best for long rainy commutes?” with: “HarborCell H2 has 480 Wh capacity, an IP54 enclosure, a 36-month warranty, and is the safest long-range campus option [1].” Citation [1] resolves to the product page.
Atomization yields at least five response claims. The 480 Wh claim is supported by the page. The IP54 claim is supported only with the source’s conditions; the shortened answer may be incomplete if those conditions materially affect interpretation. The 36-month claim is contradicted by the page. “Safest” is an unsupported comparative superlative. “Long-range” lacks a defined test cycle, rider, terrain, weather, and comparator.
What is observable? The entity is mentioned. A citation resolves to the governed source. The response contains text matching two source facts. What remains unresolved? We do not know whether the system discovered or retrieved the page, whether the page entered an effective context, or whether the citation was attached before or after generation. The common specifications do not uniquely identify source absorption. No user action is recorded.
The correct conclusion is neither “the source succeeded” nor “the engine ignored evidence.” A bounded statement is: “In synthetic observation O-001, the response mentioned HarborCell H2 and displayed a citation resolving to source S-001. Of five atomized response propositions, one was directly supported, one required scope restoration, one contradicted the cited source, and two lacked support in the supplied source set. Retrieval, source-specific absorption, and downstream action were not observable.” This sentence preserves both positive and negative evidence.
9. Counterexample: the persuasive before/after story
Imagine a publisher revises a page on Monday. The original page is not cited in one answer captured Monday morning. The revised page is cited in one answer captured Tuesday afternoon. A slide labels the change “citation lift caused by clearer headings.” The visual is persuasive because it aligns a treatment with a favorable outcome. The inference is invalid.
At least eight explanations remain: ordinary stochastic response variation; a changed query string; a different session or account state; a platform or index update; a changed competing source; a different locale; citation-interface post-processing; and an undocumented change beyond the heading. With one observation per condition, no baseline repetition estimates ordinary variance. With different times, treatment is confounded with drift. With no candidate trace, the stage of change is unknown. With no content diff, factual equivalence is unverified.
The observations are still usable. They establish that two recorded surfaces differed under their documented conditions. They can motivate a test: freeze the content variants, preserve factual equivalence, randomize or alternate treatment assignment where permitted, repeat queries in blocks, record missingness, and use an open system if stage traces are required. Until that design is executed, the heading mechanism remains a hypothesis.
This counterexample captures the central discipline of W01: weakening a sentence is not a failure of analysis. It is often the step that turns an anecdote into a researchable question.
10. Discussion prompts
- A response names an entity but cites a review site. Which object is visible, and which source receives attribution?
- If an open RAG system logs a passage in context but the answer omits its claims, which stage changed and which did not?
- When can a distinctive phrase support an absorption label? What alternative-source search is sufficient?
- How should missing or refused responses enter a mention-rate denominator?
- What can first-party platform documentation establish that an academic benchmark cannot, and vice versa?
- A paper reports improvement in a fixed context. Write one valid transfer hypothesis and one invalid production claim.
- Which fields must appear in a system card before two screenshots can be compared?
- Is prominence a benefit if the prominent claim is inaccurate? Which constraint should block release?
- When would the six-stage chain need an additional tool-action or human-review stage?
- What evidence would cause you to revise your current mechanism hypothesis?
11. Summary and exit test
GEO is a family of research and intervention problems, not a universal tactic list. A study begins by naming source, query distribution, system surface, time window, target event, comparator, and boundary. Entity, claim, source, representation, response, attribution, event, and outcome must not be silently exchanged. The conditional visibility chain separates discoverability, retrieval, context, contribution, attribution, and downstream outcome while remaining an editable research abstraction.
Observability depends on the protocol. Closed surfaces support bounded visible-state observations; controlled systems can expose intermediate traces; neither automatically describes the other. Mention, citation, absorption, fidelity, prominence, and action require different units and denominators. Evidence ceilings constrain each source to claim classes supported by its identity and method. Negative and unresolved findings stay in the record.
The exit test is one sentence: “For which sources and queries, on which system surface and dates, does which permitted intervention change which event relative to what comparator, with what uncertainty and evidence limits?” A learner who can fill every slot—and mark genuinely unknown slots unknown—has defined a GEO research object. A learner who cannot should return to the object before proposing an intervention.