# Web and Local Reference Catalog for the GEO Monograph and Course

**Catalog freeze:** 2026-08-24 (Asia/Shanghai)  
**Search mode:** ordinary web search and direct source inspection; the research-assistant workflow was explicitly not used.  
**Purpose:** a controlled source ledger for the monograph and the online course. Inclusion here is not an endorsement of every claim made by a source.

> **ID namespace:** supplied research papers use the public form
> `PAPER-01`–`PAPER-42`. Platform/protocol entries use `PLAT-01`–`PLAT-15` in
> every human-readable row. The generated machine view alone preserves the
> historical internal P-family keys for frozen route compatibility and adds an
> explicit `displayId`. The display namespaces prevent a paper and a platform
> source from sharing one visible identifier without rewriting frozen
> provenance records.

## Catalog boundary: curated freeze versus historical inventory

The **86 entries in this file are the current curated freeze, not the complete
historical research universe**. A separate reconciliation of eight legacy
course/monograph files found 1,961 Markdown-link occurrences, 680 normalized
URLs, and 200 domains. Only 27 URLs directly overlap this catalog; after the
latest 42-paper corpus and two official-document aliases are reconciled, 41
source identities are known to be represented somewhere in the current
system. The remaining material is retained—not discarded—in a supplementary
candidate queue or a legacy watch/context corpus.

See [`legacy_reference_reconciliation.md`](legacy_reference_reconciliation.md)
for the full inventory, exact links, type and domain counts, overlap method,
primary/official migration list, GitHub and media queues, white-paper ledger,
dead/recovered links, and the four-batch integration plan. Adding a source to a
future catalog version requires a new dated freeze; this 86-entry denominator
and its 12 Integrated / 62 Candidate / 12 Context-only accounting must not be
silently changed.

Ordinary-search deltas are tracked separately in
[`reference_search_delta_2026-08-24_pass2.md`](reference_search_delta_2026-08-24_pass2.md)
and
[`reference_search_delta_2026-08-24_pass3.md`](reference_search_delta_2026-08-24_pass3.md).
Pass 3 records status-verified standards, regulatory measures, and transparency
guidance; none is silently added to this frozen denominator.
Pass 4 is recorded in
[`reference_search_delta_2026-08-24_pass4.md`](reference_search_delta_2026-08-24_pass4.md)
and audits tutorials, repositories, practitioner explanations, platform-version
documentation, and white papers without changing the freeze.
Pass 5 is recorded in
[`reference_search_delta_2026-08-24_pass5.md`](reference_search_delta_2026-08-24_pass5.md)
and cross-checks paper-linked repositories, open datasets, venue records, and
unresolved artifact gaps; it also leaves the 86-entry denominator unchanged.
Pass 6 is recorded in
[`reference_search_delta_2026-08-24_pass6.md`](reference_search_delta_2026-08-24_pass6.md)
and adds attribution/citation evaluation, reproducibility, conflicting evidence,
current platform/protocol records, recommendation studies, and two university
course comparators without changing the freeze. Pass 7 is recorded in
[`reference_search_delta_2026-08-24_pass7.md`](reference_search_delta_2026-08-24_pass7.md)
and separates ACL 2026 publication-status upgrades, additional course-design
comparators, discovery indexes, commercial curricula, and domestic white-paper
red-team fixtures; it likewise preserves the 42-paper and 86-source denominators.

## 1. How to read this catalog

### Status

- **Integrated** — a bounded idea, structure, local artifact, or design principle is already represented in the current project architecture. It does **not** mean that every statement in the source has been adopted.
- **Candidate** — verified enough to support a planned chapter, lesson, figure, or lab, but not yet integrated claim by claim.
- **Context-only** — useful for horizon scanning, examples, or hypotheses. It must not be used as the sole authority for a scientific, legal, standards, or cross-platform claim.

### Evidence ceilings

- **L — Local controlled artifact.** Supports provenance, corpus membership, and project decisions only; paper conclusions still require paper-level reading.
- **N — Normative or official rule.** Supports requirements within its stated jurisdiction or protocol scope; it does not prove ranking or citation effects.
- **A — Primary research or benchmark.** Supports claims bounded by the published design, data, models, and date; no automatic generalization to current closed platforms.
- **B — Official academic course, textbook, or maintained research implementation.** Supports pedagogy, definitions, and reproducible implementation patterns; not a substitute for empirical evidence about closed engines.
- **C — First-party platform documentation.** Authoritative about the named platform at the stated date only; not cross-engine evidence and not a visibility guarantee.
- **D — Vendor observational or quasi-experimental study.** Useful descriptive evidence when methods are inspectable; proprietary sampling, parsers, and commercial incentives cap causal inference.
- **E — Practitioner or secondary synthesis.** Useful for terminology, checklists, and testable hypotheses; technical claims must be checked against primary papers, official documentation, or controlled experiments.

### Module map

| Code | Monograph module | Code | Course module / lab |
|---|---|---|---|
| B1 | Field definition, history, actors, and research questions | C1 | Foundations and end-to-end system map |
| B2 | Query transformation, retrieval, ranking, and generation | C2 | Technical discoverability and crawler audit |
| B3 | Evidence, attribution, citation, and provenance | C3 | Retrieval, reranking, and corpus lab |
| B4 | Metrics, benchmarks, uncertainty, and causal design | C4 | Evidence and citation diagnostics lab |
| B5 | Content, page, entity, and site interventions | C5 | Repeated-run visibility measurement lab |
| B6 | Multimodal, multilingual, and agentic systems | C6 | Controlled intervention and ablation lab |
| B7 | Manipulation, safety, fairness, and governance | C7 | Adversarial testing and governance clinic |
| B8 | Deployment, standards, accessibility, and operations | C8 | Reproducible capstone and evidence dossier |

The **Verified fact → proposed use** column deliberately separates what was checked from what remains an editorial proposal.

## 2. Existing and project-local anchors

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| L01 | [Latest Zotero Paper Export (42 Records)](/readings/papers) | Local CSV corpus | 2026-08-24 | 2026-08-24 | **Verified:** 42 records, 87 fields, 42 DOI values, 41 URLs, 41 abstracts, and 42 local PDFs; SHA-256 `c149a69140b0dcedc1855f8cbb02fc5cb1e1346470306c76cd68fbafc2e50468`. **Proposed:** canonical reading-list input and claim-level paper queue. | L — proves local corpus state, not the validity of paper claims. | B1–B7; C1, C4–C8 | **Integrated** |
| L02 | [Normalized GEO Paper Catalog](/readings/papers) | Local normalized JSON | 2026-08-24 | 2026-08-24 | **Verified:** machine-readable normalization of the latest CSV. **Proposed:** generate reference views, tags, module links, and version checks from one structured source. | L — editorial routing only until each paper is read and verified. | B1–B7; C8 | **Integrated** |
| L03 | [Latest Paper Audit and Module Mapping](/provenance) | Local audit | 2026-08-24; route control updated 2026-08-25 | 2026-08-25 | **Verified:** preserves the initial 42-record triage and five metadata issues. The separate machine-readable `sources/paper_route_crosswalk_v1.6.json` is authoritative for final 28-chapter monograph, actual 12-chapter Core Notes citations, 16-week course, and evidence-card claim routes: 42 papers, 48 paper–claim edges, 17 unique claims, and three zero-edge cards. **Proposed:** retain the Markdown table as historical triage; use the crosswalk for v1.6 route checks. | L — routing is an editorial control, not human scientific approval or independent reproduction. | B1–B7; C1–C8 | **Integrated** |
| L04 | [An Introduction to Flow Matching and Diffusion Models](https://arxiv.org/abs/2506.02070) | arXiv tutorial / lecture notes | Submitted 2025-06-02; v3 2026-03-18 | 2026-08-24 | **Verified:** tutorial-style notes with a compact title block, parts, appendices, equations, figures, and pedagogical environments. **Proposed:** emulate information architecture and teaching-role differentiation, not subject matter. | B — design/pedagogy reference; the arXiv record is not GEO evidence. | B8; C8 | **Integrated** |
| L05 | [Local TeX Source Archive for arXiv:2506.02070v3](/provenance#treatment) | Local TeX source reference | v3, 2026-03-18 | 2026-08-24 | **Verified:** 48 safe archive members; modular `main.tex`, seven parts, five appendices; SHA-256 `b69a997d626dcbb94e3b91d83e3077251779f2d9ab767c4ea4209d84ea135fbf`. The declared TeX Live 2025/`pdflatex` path reproduces an 84-page PDF; six page roles were visually sampled. The build remains warning-bearing, untagged, and has blank Title/Author metadata. License is CC BY-NC-ND 4.0. **Proposed:** transfer only abstract design principles via original prose and original figures, while enforcing stricter GEO release gates. | L/B — architecture reference only; reproducible does not mean publication-ready; no text or figure adaptation. | B8; C8 | **Integrated** |
| L06 | [Stanford CS336: Language Modeling from Scratch](https://cs336.stanford.edu/) | Official university course | Spring 2026 | 2026-08-24 | **Verified:** a schedule-first course page with lectures, assignments, readings, and recordings. **Proposed:** use its high-density weekly rhythm, assignment visibility, and prerequisite signaling. | B — strong curriculum model; not GEO-specific. | B2, B8; C1, C3, C8 | **Integrated** |
| L07 | [Stanford CS229: Machine Learning — Main Notes](https://cs229.stanford.edu/main_notes.pdf) | Official university notes | 2026-08-23 edition | 2026-08-24 | **Verified:** a 278-page coherent mathematical notes volume. **Proposed:** benchmark notation discipline, prerequisite refreshers, exercises, and appendix structure. | B — mathematical pedagogy; not evidence about generative search. | B4, B8; C1, C5, C8 | **Integrated** |
| L08 | [cuhkgeo.com self-description](https://www.cuhkgeo.com/) | Subject-published context page | Current 2026 snapshot | 2026-08-25 | **Verified:** the page presents research themes including GEO, agents/online decision-making, trusted AI, and control/optimization. **Proposed:** treat those claims as dated research-context questions and use restrained visual cues only. | B/E — self-description only; it does not verify this course, institutional ownership, staffing, affiliation, partnership, sponsorship, endorsement, or individual scientific claims. | B1, B6, B7; C1, C7 | **Integrated** |
| L09 | [FrontMind](https://www.frontmind.net/) | Industry website | Current 2026 snapshot | 2026-08-24 | **Verified:** first-party industry framing organized around understanding, growth, and embedding. **Proposed:** retain the independently rewritten research loop “observe → explain → intervene → validate” and use brand cues sparingly. | E — industry framing/design only; no scientific or performance claims. | B1, B4, B5; C1, C5, C6 | **Integrated** |

## 3. Primary research, academic tutorials, courses, and benchmark programs

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| A01 | [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735) | Primary paper; KDD 2024 | Submitted 2023-11-16; v3 2024-06-28 | 2026-08-24 | **Verified:** introduces GEO, GEO-Bench, visibility metrics, and controlled content interventions; reports gains up to 40% in its experimental setting. **Proposed:** field origin, baseline methods, and a reproduction/critique lab. | A — peer-reviewed, but results are setup-, engine-, metric-, and date-bounded. | B1, B4, B5; C1, C5, C6 | **Integrated** |
| A02 | [KDD 2024 Research Track Papers](https://www.kdd.org/kdd2024/research-track-papers/) | Official venue index | 2024 | 2026-08-24 | **Verified:** official venue listing for the GEO paper. **Proposed:** use only to verify publication status and venue metadata. | A/B — bibliographic verification, not independent replication. | B1; C1 | **Integrated** |
| A03 | [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks](https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html) | NeurIPS primary paper | 2020 | 2026-08-24 | **Verified:** formalizes a parametric-plus-non-parametric retrieval/generation architecture and evaluates knowledge-intensive tasks. **Proposed:** canonical mechanism figure and notation for the retrieval-to-generation pipeline. | A — foundational open-system evidence; not a disclosure of current commercial search stacks. | B2, B3; C1, C3, C4 | **Candidate** |
| A04 | [Dense Passage Retrieval for Open-Domain Question Answering](https://aclanthology.org/2020.emnlp-main.550/) | EMNLP primary paper | 2020 | 2026-08-24 | **Verified:** evaluates dual-encoder dense retrieval for open-domain QA. **Proposed:** contrast sparse, dense, and hybrid candidate retrieval in the retrieval lab. | A — benchmark-bounded; later embedding models require fresh evaluation. | B2, B4; C3 | **Candidate** |
| A05 | [Enabling Large Language Models to Generate Text with Citations](https://aclanthology.org/2023.emnlp-main.398/) | EMNLP primary paper / ALCE benchmark | 2023 | 2026-08-24 | **Verified:** evaluates correctness and completeness of citations for long-form generation. **Proposed:** distinguish answer quality, citation correctness, citation completeness, and source quality. | A — benchmark and model snapshot bounded. | B3, B4; C4, C5 | **Candidate** |
| A06 | [Text REtrieval Conference (TREC)](https://trec.nist.gov/) | NIST evaluation program | Ongoing; 2026 program visible | 2026-08-24 | **Verified:** long-running shared evaluation program; current tracks include retrieval-augmented generation. **Proposed:** model the course’s evaluation contracts, run files, qrels, and reproducibility package on TREC conventions. | A/B — authoritative benchmark program; each track’s task and judgments define its ceiling. | B2–B4; C3–C5, C8 | **Candidate** |
| A07 | [TREC 2024 Retrieval-Augmented Generation Track Data](https://trec.nist.gov/data/rag2024.html) | Official benchmark data page | Created 2025-03-11 for TREC 2024 | 2026-08-24 | **Verified:** publishes topics, qrels, nuggets, and citation-support judgments. **Proposed:** adapt a small lawful subset or analogous original set for answer/citation scoring labs. | A — benchmark evidence only; licensing and redistribution must be checked per resource. | B3, B4; C4, C5, C8 | **Candidate** |
| A08 | [Overview of the TREC 2025 Retrieval-Augmented Generation Track](https://trec.nist.gov/pubs/trec34/papers/Overview_rag.pdf) | Official track overview | TREC 2025 / published 2026 | 2026-08-24 | **Verified:** documents sentence-level answer evaluation and citation-oriented judgments. **Proposed:** update the rubric beyond citation counts to claim-level support. | A — task-specific evaluation design, not a universal quality metric. | B3, B4; C4, C5 | **Candidate** |
| A09 | [Stanford CS276: Information Retrieval and Web Search](https://web.stanford.edu/class/cs276/index.html) | Official university course archive | Spring 2019 | 2026-08-24 | **Verified:** covers indexing, retrieval, BM25, evaluation/NDCG, crawling, and web search. **Proposed:** prerequisite bridge and classical-IR comparison boxes. | B — durable foundations; web-platform details are dated. | B2, B4, B8; C1–C3 | **Candidate** |
| A10 | [Introduction to Information Retrieval](https://nlp.stanford.edu/IR-book/) | Academic textbook, free online edition | Book 2008; online 2009 | 2026-08-24 | **Verified:** canonical treatment of indexing, scoring, evaluation, crawling, and link analysis. **Proposed:** prerequisite appendix and notation source. | B — foundational, not current neural/generative-search evidence. | B2, B4, B8; C1, C3 | **Candidate** |
| A11 | [CMU 11-442/11-642: Search Engines](https://boston.lti.cs.cmu.edu/classes/11-442/index.html) | Official university course | Fall 2026; page updated 2026-05-28 | 2026-08-24 | **Verified:** covers text search, RAG, and recommender/search systems in one contemporary course. **Proposed:** benchmark sequencing from classical retrieval to generative applications. | B — current academic syllabus; individual lecture claims need readings. | B2–B4; C1, C3–C5 | **Candidate** |
| A12 | [Next Generation Search: LLM, Conversational AI and Query Prediction](https://www.oulu.fi/en/education-search/next-generation-search-llm-conversational-ai-and-query-prediction) | Official doctoral course | 2025–2026 | 2026-08-24 | **Verified:** combines LLM-based IR, vector search, reranking, conversational search, and query prediction. **Proposed:** cross-check advanced-course outcomes and workload. | B — syllabus-level evidence only. | B2, B6; C1, C3, C8 | **Candidate** |
| A13 | [Information Retrieval and Search Engines](https://onderwijsaanbod.kuleuven.be/2025/syllabi/e/H02C8BE.htm) | Official university syllabus | 2025–2026 | 2026-08-24 | **Verified:** integrates IR/search with LLM-based conversational agents. **Proposed:** compare European learning outcomes and assessment balance. | B — syllabus-level evidence only. | B2, B6; C1, C3 | **Candidate** |
| A14 | [Stanford CS224V: Conversational Virtual Assistants with Deep Learning](https://web.stanford.edu/class/cs224v/) | Official project-oriented university course | Fall 2025 | 2026-08-24 | **Verified:** explicitly asks how to perform RAG without hallucination and retrieve from databases/knowledge graphs, with final projects. **Proposed:** inspiration for the course capstone and evidence-curation studio. | B — pedagogical model; not a guarantee that any method eliminates hallucination. | B2, B3, B6; C3, C4, C8 | **Candidate** |
| A15 | [Recent Advances in Generative Information Retrieval](https://generative-ir.github.io/) | SIGIR 2024 tutorial site | 2024-07-14 | 2026-08-24 | **Verified:** tutorial materials and bibliography on generative retrieval. **Proposed:** scholarly bridge from document retrieval to generative IR; use slides only within their license. | B — expert tutorial synthesis; primary papers remain claim authorities. | B2, B4; C1, C3 | **Candidate** |
| A16 | [Full Stack LLM Bootcamp](https://fullstackdeeplearning.com/llm-bootcamp/) | Open practical course | 2023 | 2026-08-24 | **Verified:** covers augmented language models, deployment, UX, and LLM operations. **Proposed:** derive production checklists and end-to-end project rhythm. | B — practical engineering reference; framework and API details age quickly. | B6, B8; C3, C8 | **Candidate** |
| A17 | [Hugging Face Agents Course: Agentic RAG](https://huggingface.co/learn/agents-course/en/unit2/smolagents/retrieval_agents) | Official interactive tutorial / Colab | Current 2026 snapshot | 2026-08-24 | **Verified:** demonstrates query reformulation, decomposition, retrieval, reranking, and multi-step validation. **Proposed:** optional agentic-retrieval lab with frozen dependencies. | B — runnable tutorial; library behavior and examples can drift. | B2, B6; C3, C8 | **Candidate** |
| A18 | [Building and Evaluating Advanced RAG Applications](https://www.deeplearning.ai/courses/building-evaluating-advanced-rag/) | Vendor-hosted educational short course | 2023-11 | 2026-08-24 | **Verified:** approximately two hours with videos and code on retrieval evaluation and advanced chunk/retrieval strategies. **Proposed:** use as optional preparatory practice, not as the course’s evidence authority. | E/B — instructional value; named metrics and product concepts need independent validation. | B2–B4; C3, C4 | **Candidate** |

## 4. GitHub repositories and reproducible implementation references

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| G01 | [GEO: Official Code and GEO-Bench](https://github.com/GEO-optim/GEO) | Official research repository; Apache-2.0 | 2024 snapshot | 2026-08-24 | **Verified:** paper-linked code, benchmark assets, and evaluation implementation. **Proposed:** pin a commit, reproduce one bounded baseline, then document deviations and failures. | B/A — strongest implementation companion to A01; code does not independently validate results. | B4, B5; C5, C6, C8 | **Integrated** |
| G02 | [BEIR: Heterogeneous Benchmark for Information Retrieval](https://github.com/beir-cellar/beir) | Official research repository / benchmark | NeurIPS 2021; release 2.2.0 on 2025-06-04 | 2026-08-24 | **Verified:** heterogeneous IR evaluation across multiple datasets and retrieval systems. **Proposed:** small sparse-vs-dense-vs-hybrid lab and domain-shift discussion. | B/A — reproducible benchmark; dataset licenses and domain fit vary. | B2, B4; C3, C5, C8 | **Candidate** |
| G03 | [Pyserini](https://github.com/castorini/pyserini) | Maintained academic IR toolkit | Current 2026 snapshot | 2026-08-24 | **Verified:** reproducible sparse, dense, and hybrid first-stage retrieval over Lucene/Faiss with prebuilt indexes and qrels. **Proposed:** default retrieval lab toolkit with a pinned environment. | B — implementation authority for the toolkit, not evidence about closed platforms. | B2, B4; C3, C8 | **Candidate** |
| G04 | [RankLLM](https://github.com/castorini/rank_llm) | Academic LLM-reranking toolkit | Current 2026 snapshot | 2026-08-24 | **Verified:** supports reproducible pointwise/listwise LLM reranking workflows. **Proposed:** compare BM25, dense, and LLM reranking under fixed candidate sets and budgets. | B — toolkit and associated-paper evidence only; model/API drift must be frozen. | B2, B4; C3, C5 | **Candidate** |
| G05 | [Massive Text Embedding Benchmark (MTEB)](https://github.com/embeddings-benchmark/mteb) | Open benchmark framework | 2022–current | 2026-08-24 | **Verified:** standardized evaluation across embedding tasks including retrieval. **Proposed:** teach model selection as dataset- and language-dependent rather than leaderboard universalism. | B/A — benchmark result ceiling follows dataset coverage and evaluation protocol. | B2, B4, B6; C3, C5 | **Candidate** |
| G06 | [Ragas](https://github.com/vibrantlabsai/ragas) | Open-source RAG evaluation framework | Current 2026 snapshot | 2026-08-24 | **Verified:** implements dataset generation, RAG evaluation, and feedback workflows. **Proposed:** optional lab scaffold while separately validating metric definitions and judge reliability. | B — framework metrics are operationalizations, not ground truth. | B3, B4; C4, C5, C8 | **Candidate** |
| G07 | [DSPy](https://github.com/stanfordnlp/dspy) | Academic LM-programming and optimization framework | ICLR 2024 / current | 2026-08-24 | **Verified:** modular LM programs and metric-driven compilation/optimization. **Proposed:** controlled multi-objective intervention lab with train/dev/test separation. | B/A — framework is credible; outcomes depend on metric, data, and model. | B4–B6; C6, C8 | **Candidate** |
| G08 | [Promptfoo](https://github.com/promptfoo/promptfoo) | Open-source evaluation and red-team toolkit | Current 2026 snapshot | 2026-08-24 | **Verified:** supports test matrices, assertions, red-team checks, and CI workflows. **Proposed:** capstone regression suite for prompts, answers, citations, and policy checks. | B — tooling reference; generated scores require audited assertions and judges. | B4, B7, B8; C5, C7, C8 | **Candidate** |
| G09 | [TruLens](https://github.com/truera/trulens) | Open-source tracing/evaluation toolkit | Current 2026 snapshot | 2026-08-24 | **Verified:** supports tracing and feedback functions for retrieval and generation. **Proposed:** optional observability comparison with Ragas; do not import vendor benchmark claims. | B — implementation only; metric validity must be independently assessed. | B3, B4, B8; C4, C5, C8 | **Candidate** |
| G10 | [The llms.txt Proposal Repository](https://github.com/AnswerDotAI/llms-txt) | Community proposal / repository | Proposed 2024-09-03; updated 2026 | 2026-08-24 | **Verified:** a voluntary proposal for an `/llms.txt` file, not an IETF/W3C/ISO web standard. Google’s 2026 guide explicitly says Google Search ignores it. **Proposed:** use only in an experiment contrasting proposals with verified platform behavior. | E — no general ranking/citation claim and no cross-platform requirement. | B8; C2, C6 | **Context-only** |

## 5. First-party platform guidance and open-web protocols

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| PLAT-01 | [Google’s Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) | Official Google Search documentation | Last updated 2026-07-10 | 2026-08-24 | **Verified:** Google says existing SEO fundamentals apply; describes RAG and query fan-out; states indexing/serving are not guaranteed; says Google Search ignores `llms.txt` and no special AI markup is required. **Proposed:** primary platform-specific boundary text and myth-busting lab. | C — authoritative for Google Search only; no cross-engine or outcome guarantee. | B1, B2, B5, B8; C1, C2, C6 | **Candidate** |
| PLAT-02 | [Google Search Guidance on Using Generative AI Content](https://developers.google.com/search/docs/fundamentals/using-gen-ai-content) | Official Google Search documentation | Current 2026 snapshot | 2026-08-24 | **Verified:** mass generation without user value may violate scaled-content-abuse policy; emphasizes accuracy, quality, relevance, and appropriate creation context. **Proposed:** responsible publishing checklist. | C — Google-specific policy guidance, not a scientific quality metric. | B5, B7, B8; C6, C7 | **Candidate** |
| PLAT-03 | [Google General Structured Data Guidelines](https://developers.google.com/search/docs/appearance/structured-data/sd-policies) | Official Google Search documentation | Last updated 2026-07-10 | 2026-08-24 | **Verified:** valid structured data creates eligibility for supported features but does not guarantee display. **Proposed:** schema validation lab with explicit “eligibility ≠ citation causality” language. | C — Google eligibility/policy only. | B5, B8; C2, C6 | **Candidate** |
| PLAT-04 | [Introducing AI Performance in Bing Webmaster Tools Public Preview](https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview) | Official Microsoft/Bing product guidance | 2026-02-10 | 2026-08-24 | **Verified:** preview reports citations, cited pages, sampled grounding queries, and trends across specified Microsoft surfaces; counts do not indicate rank, authority, or placement. **Proposed:** measurement-interface case study and metric-definition critique. | C — early-preview, aggregated, Microsoft-surface-specific data. | B4, B8; C2, C5 | **Candidate** |
| PLAT-05 | [Keeping Content Discoverable with Sitemaps in AI-Powered Search](https://blogs.bing.com/webmaster/July-2025/Keeping-Content-Discoverable-with-Sitemaps-in-AI-Powered-Search) | Official Microsoft/Bing guidance | 2025-07-31 | 2026-08-24 | **Verified:** recommends accurate XML sitemaps and `lastmod`; explicitly says no tool guarantees appearance in AI-generated results. **Proposed:** sitemap freshness exercise. | C — Bing-specific operational guidance, not a citation guarantee. | B8; C2 | **Candidate** |
| PLAT-06 | [Bing Support for the `data-nosnippet` HTML Attribute](https://blogs.bing.com/webmaster/October-2025/Bing-Introduces-Support-for-the-data-nosnippet-HTML-Attribute) | Official Microsoft/Bing guidance | 2025-10-15 | 2026-08-24 | **Verified:** marked content may remain indexed while being excluded from Bing snippets and AI summaries. **Proposed:** publisher-control lab contrasting crawl, index, retrieval, and display controls. | C — Bing-supported environments only. | B3, B8; C2, C4 | **Candidate** |
| PLAT-07 | [Publishers and Developers FAQ](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq) | Official OpenAI help documentation | Page reported updated 2026-07-31 | 2026-08-24 | **Verified:** distinguishes OAI-SearchBot access for search inclusion from GPTBot training controls and documents ChatGPT referral UTM parameters. **Proposed:** crawler-policy matrix and log-analysis lab. | C — OpenAI-product-specific and changeable; recheck before publication. | B2, B8; C2, C5 | **Candidate** |
| PLAT-08 | [Searching the Web with ChatGPT](https://help.openai.com/en/articles/9237897/chatgpt-search) | Official OpenAI help documentation | Page reported updated 2026-08-22 | 2026-08-24 | **Verified:** describes current web-search behavior, source review, query rewriting, and site eligibility; explicitly says placement is not guaranteed. **Proposed:** UI/surface observation protocol, with screenshots dated and labeled. | C — product behavior can change; not an algorithm specification. | B1, B3, B4; C1, C4, C5 | **Candidate** |
| PLAT-09 | [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) | Official Perplexity documentation | Current 2026 snapshot | 2026-08-24 | **Verified:** distinguishes PerplexityBot and user-triggered retrieval behavior and publishes crawler information. **Proposed:** first-party source for crawler identification and access-log parsing. | C — Perplexity-specific; user agents and IP ranges must be revalidated at lab time. | B2, B8; C2 | **Candidate** |
| PLAT-10 | [How Perplexity Follows robots.txt](https://www.perplexity.ai/help-center/en/articles/10354969-how-does-perplexity-follow-robots-txt) | Official Perplexity help documentation | Updated 2026-07-16 | 2026-08-24 | **Verified:** explains Perplexity’s stated treatment of robots controls. **Proposed:** compare declared policy with server-log observations without inferring intent from a user agent alone. | C — stated platform policy, not independent compliance evidence. | B7, B8; C2, C7 | **Candidate** |
| PLAT-11 | [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html) | IETF Standards Track RFC | 2022-09 | 2026-08-24 | **Verified:** normative syntax and matching rules for robots.txt. **Proposed:** authoritative protocol basis for all crawler-control examples. | N — protocol semantics only; REP is not access authorization and does not guarantee crawler compliance or visibility. | B8; C2, C7 | **Candidate** |
| PLAT-12 | [Sitemaps XML Protocol](https://www.sitemaps.org/protocol.html) | Open web protocol documentation | Current protocol page; n.d. | 2026-08-24 | **Verified:** defines sitemap structure and limits. **Proposed:** technical discoverability lab and validation checklist. | N/B — protocol syntax only; submission does not guarantee crawling, indexing, ranking, or citation. | B8; C2 | **Candidate** |
| PLAT-13 | [Schema.org for Developers](https://schema.org/docs/developers.html) | Community vocabulary documentation | Version 30.0, 2026-03-19 | 2026-08-24 | **Verified:** documents a shared structured-data vocabulary. **Proposed:** semantic entity/page representation exercise. | N/B — vocabulary semantics only; no platform visibility effect is implied. | B5, B8; C2, C6 | **Candidate** |
| PLAT-14 | [IndexNow Protocol Documentation](https://www.indexnow.org/documentation) | Open URL-notification protocol | Current 2026 snapshot | 2026-08-24 | **Verified:** specifies how sites notify participating engines of added, updated, or deleted URLs. **Proposed:** instrument notification latency and downstream crawl observations. | N/B — notification, not a promise of crawl, index, rank, or citation. | B8; C2, C5 | **Candidate** |
| PLAT-15 | [Web Content Accessibility Guidelines (WCAG) 2.2](https://www.w3.org/TR/WCAG22/) | W3C Recommendation | 2023-10-05; updated 2024-12-12 | 2026-08-24 | **Verified:** normative accessibility success criteria. **Proposed:** mandatory quality gate for the course site, figures, forms, and labs. | N — accessibility conformance only; no GEO performance claim. | B8; C2, C8 | **Candidate** |

## 6. Standards, official governance, and provenance guidance

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| S01 | [NIST AI Risk Management Framework 1.0](https://www.nist.gov/itl/ai-risk-management-framework) | U.S. government voluntary framework | 2023-01-26; revision in progress in 2026 | 2026-08-24 | **Verified:** voluntary framework for incorporating trustworthiness into AI design, development, use, and evaluation. **Proposed:** structure governance around Govern, Map, Measure, and Manage while flagging the active revision. | N — voluntary risk framework, not law or certification and not GEO-specific. | B7, B8; C7, C8 | **Candidate** |
| S02 | [NIST AI 600-1: Generative AI Profile](https://doi.org/10.6028/NIST.AI.600-1) | U.S. government technical report | 2024-07-26; page updated 2026-04-08 | 2026-08-24 | **Verified:** cross-sector companion profile for GenAI risks and actions. **Proposed:** risk taxonomy for evidence integrity, information security, content provenance, and incident response. | N — voluntary cross-sector guidance; implementations require contextual tailoring. | B7, B8; C7, C8 | **Candidate** |
| S03 | [ISO/IEC 42001:2023 — Artificial Intelligence Management System](https://www.iso.org/standard/42001) | International standard | 2023-12 | 2026-08-24 | **Verified:** specifies requirements for establishing, implementing, maintaining, and continually improving an AI management system. **Proposed:** organizational governance checklist and capstone audit trail. | N — official metadata/overview is public; full requirements are paywalled and must not be inferred from the overview. | B7, B8; C7, C8 | **Candidate** |
| S04 | [C2PA Content Credentials Technical Specification 2.4](https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html) | Industry technical specification | 2.4, 2026-04 | 2026-08-24 | **Verified:** specifies cryptographically bound, tamper-evident provenance manifests; expressly does not judge whether provenance is “good” or “bad.” **Proposed:** multimodal provenance and disclosure lab. | N/B — provenance integrity and validation only; not truth verification, authorship proof, or ranking evidence. | B3, B6–B8; C4, C7, C8 | **Candidate** |
| S05 | [Interim Measures for the Management of Generative AI Services](https://www.gov.cn/zhengce/zhengceku/202307/content_6891752.htm) | PRC regulation / official government text | Issued 2023-07-10; effective 2023-08-15 | 2026-08-24 | **Verified:** applies to specified public-facing generative AI services in mainland China and sets provider obligations. **Proposed:** jurisdiction-aware governance unit; obtain legal review for operational advice. | N — official legal text; not legal advice and scope/application require counsel. | B7, B8; C7 | **Candidate** |
| S06 | [Measures for Labeling AI-Generated and Synthetic Content](https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm) | PRC regulatory measure / official text | Issued 2025-03-07; effective 2025-09-01 | 2026-08-24 | **Verified:** requires specified explicit and implicit labels and related platform/provider duties. **Proposed:** publishing and provenance compliance checklist. | N — official legal measure; not legal advice. | B3, B6–B8; C7, C8 | **Candidate** |
| S07 | [GB 45438-2025 — Labeling Method for Content Generated by Artificial Intelligence](https://openstd.samr.gov.cn/bzgk/std/newGbInfo?hcno=F32EA2A561F1886CD8D606513512D547&refer=outter) | Mandatory Chinese national standard | Published 2025-02-28; effective 2025-09-01 | 2026-08-24 | **Verified:** current mandatory standard for AI-generated/synthetic content labeling methods. **Proposed:** standards crosswalk with S06 and C2PA, carefully separating mandatory labeling from optional provenance technologies. | N — normative within scope; no relevance to ranking/citation performance. | B3, B6–B8; C7, C8 | **Candidate** |
| S08 | [GB/T 45654-2025 — Basic Security Requirements for Generative AI Services](https://std.samr.gov.cn/gb/search/gbDetailed?id=33D40F1160BF5D92E06397BE0A0A5B93) | Recommended Chinese national standard | Published 2025-04-25; effective 2025-11-01 | 2026-08-24 | **Verified:** current recommended security standard for generative AI services. **Proposed:** security-governance matrix for a GEO measurement or optimization service. | N — recommended standard, not law and not a GEO-performance standard. | B7, B8; C7, C8 | **Candidate** |
| S09 | [GB/T 45652-2025 — Security Specification for Generative AI Pre-training and Fine-tuning Data](https://std.samr.gov.cn/gb/search/gbDetailedCNF?id=33D40F1160BE5D92E06397BE0A0A5B93) | Recommended Chinese national standard | Published 2025-04-25; effective 2025-11-01 | 2026-08-24 | **Verified:** current recommended standard covering pre-training and fine-tuning data security. **Proposed:** data lineage, provenance, licensing, and contamination checklist. | N — data-security scope only; not directly a content-ranking rule. | B7, B8; C7, C8 | **Candidate** |
| S10 | [GB/T 35273-2020 — Personal Information Security Specification](https://std.samr.gov.cn/gb/search/gbDetailed?id=A0280129495AEBB4E05397BE0A0AB6FE) | Recommended Chinese national standard | 2020-03-06; effective 2020-10-01; revision in progress in 2026 | 2026-08-24 | **Verified:** current personal-information security specification, with a revision process underway. **Proposed:** privacy controls for query logs, user studies, analytics, and model outputs. | N — recommended standard and subject to revision; obtain legal/privacy review. | B4, B7, B8; C5, C7, C8 | **Candidate** |
| S11 | [Project 20252040-Z-469 — General Technical Requirements of Retrieval-Augmented Generation](https://std.samr.gov.cn/gb/search/gbDetailed?id=37FC03D2E1436322E06397BE0A0AA17F) | Chinese standardization guidance project | Registered 2025-06-20; **awaiting approval** on 2026-08-24 | 2026-08-24 | **Verified:** an approval-stage project, not a published national standard. **Proposed:** standards watchlist only; refresh status before publication. | N-watch — cannot be cited as an in-force standard or requirement. | B2, B8; C3, C8 | **Context-only** |
| S12 | [Project 20254292-Z-469 — Knowledge Graph and Large Pre-trained Model Integration, Part 2: Graph RAG](https://std.samr.gov.cn/gb/search/gbDetailed?id=3BD752A86E80AF43E06397BE0A0ACD95) | Chinese standardization project | 2025 project | 2026-08-24 | **Verified:** registered project concerning Graph RAG, not a published standard. **Proposed:** horizon note in the agentic/graph retrieval chapter. | N-watch — project metadata only. | B2, B6, B8; C3, C8 | **Context-only** |

## 7. Practitioner blogs and community tutorials

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| R01 | [Generative Engine Optimization: Best Resources, Courses and Tutorials (2026)](https://thegeocommunity.com/generative-engine-optimization-resources-courses-tutorials/) | Practitioner resource map | 2026-04-24 | 2026-08-24 | **Verified:** role-based map of papers, courses, tools, and implementation topics. **Proposed:** discovery index and competitor-curriculum checklist; verify every downstream claim at its original source. | E — curated secondary source with some broad or time-sensitive claims. | B1, B8; C1, C8 | **Candidate** |
| R02 | [GEO vs SEO: How the User Funnel Has Changed](https://thegeocommunity.com/blogs/generative-engine-optimization/geo-vs-seo-user-funnel/) | Practitioner explainer | 2026-01-18 | 2026-08-24 | **Verified:** offers a conceptual contrast between link-oriented and synthesized-answer journeys. **Proposed:** introductory discussion prompt, paired with Google’s first-party guide and measurement evidence. | E — conceptual framing only; funnel claims need independent data. | B1, B4; C1, C5 | **Candidate** |
| R03 | [The Original GEO Paper Explained](https://thegeocommunity.com/blogs/generative-engine-optimization/geo-princeton-paper-original-study/) | Secondary paper explainer | 2026-02-17 | 2026-08-24 | **Verified:** practitioner-oriented explanation of A01. **Proposed:** optional pre-reading for nontechnical learners; all numerical claims cite A01 instead. | E — secondary explanation, never the primary citation for paper results. | B1, B5; C1, C6 | **Candidate** |
| R04 | [robots.txt for AI Bots: What to Allow, What to Block, and Why](https://thegeocommunity.com/blogs/generative-engine-optimization/robots-txt-ai-bots/) | Practitioner implementation tutorial | 2026-02-09 | 2026-08-24 | **Verified:** provides configuration examples and crawler taxonomy. **Proposed:** debugging cases only after rechecking each user agent against vendor docs and RFC 9309. | E — bot identities, purposes, and policies are time-sensitive; no compliance assumption. | B8; C2, C7 | **Candidate** |
| R05 | [Why LLM Answers Do Not Show Citations: 20 Reasons](https://thegeocommunity.com/blogs/generative-engine-optimization/why-llm-answers-dont-show-citations/) | Practitioner diagnostic taxonomy | 2026-06-23 | 2026-08-24 | **Verified:** separates retrieval, generation, citation scoring, post-processing, and UI failure hypotheses. **Proposed:** convert into a falsifiable diagnostic worksheet anchored to ALCE, TREC RAG, and controlled logs. | E — useful hypotheses, not evidence that hidden platform stages work as described. | B2–B4; C4, C5 | **Candidate** |
| R06 | [Log File Analysis for AI Bots](https://thegeocommunity.com/blogs/generative-engine-optimization/log-file-analysis-ai-bots-geo/) | Practitioner technical tutorial | 2026 | 2026-08-24 | **Verified:** presents server-log parsing as an observation method. **Proposed:** privacy-safe lab distinguishing claimed bot identity, verified IP, request, response, and downstream outcome. | E — method inspiration; identity verification must use first-party data and reverse/forward DNS where applicable. | B4, B8; C2, C5 | **Candidate** |
| R07 | [Visibility Has Two Axes: Repeated-Sampling Variance and Drift](https://thegeocommunity.com/blogs/generative-engine-optimization/visibility-two-axes-repeated-sampling-variance-drift/) | Practitioner research synthesis | 2026-08-18 | 2026-08-24 | **Verified:** proposes separating within-time repeated-sampling variance from between-time drift. **Proposed:** motivate repeated measures and time-blocked designs, with claims grounded in the local paper corpus. | E — derived synthesis; the decomposition must be formalized and empirically validated. | B4; C5, C6 | **Candidate** |

## 8. Vendor and industry studies / white-paper layer

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| V01 | [Insights from 55.8M AI Overviews Across 590M Searches](https://ahrefs.com/blog/insights-from-56-million-ai-overviews/) | Vendor observational study | 2025-05-19; data updated in-page | 2026-08-24 | **Verified:** reports a very large Ahrefs-owned SERP sample and discloses important coverage limitations. **Proposed:** example of scale, sampling frames, and why headline percentages need denominators and dates. | D — descriptive only; proprietary index/parser, logged-out sampling, and no causal inference. | B4; C5 | **Context-only** |
| V02 | [38% of AI Overview Citations Pull from the Top 10](https://ahrefs.com/blog/ai-overview-citations-top-10/) | Vendor observational study | 2026-03-02 | 2026-08-24 | **Verified:** analyzes 863,000 SERPs and four million AI Overview URLs with an updated parser. **Proposed:** platform-drift case study and replication-design prompt. | D — Google/Ahrefs sample only; associations do not reveal ranking mechanisms. | B2, B4; C5 | **Context-only** |
| V03 | [We Tracked 1,885 Pages Adding Schema; AI Citations Barely Moved](https://ahrefs.com/blog/schema-ai-citations/) | Vendor matched difference-in-differences study | 2026-05-11 | 2026-08-24 | **Verified:** tracks 1,885 treated and 4,000 control pages and reports no major citation uplift across tested surfaces. **Proposed:** teach correlation vs treatment effects and challenge “schema guarantees citations.” | D — stronger quasi-experimental design, but proprietary data, treatment timing uncertainty, and residual confounding remain. | B4, B5; C5, C6 | **Candidate** |
| V04 | [Why 62% of AI Citations Do Not Lead to Brand Mentions](https://www.semrush.com/blog/the-ghost-citations-study/) | Vendor observational study | 2026-06-09 | 2026-08-24 | **Verified:** distinguishes cited and mentioned outcomes using 3,981 domain appearances, 115 prompts, 14 countries, and four named surfaces. **Proposed:** metric-separation exercise, not a population estimate. | D — small prompt set and proprietary sampling; descriptive only. | B3, B4; C4, C5 | **Context-only** |
| V05 | [Semrush AI Overviews Study](https://www.semrush.com/blog/semrush-ai-overviews-study/) | Vendor longitudinal observational study | 2025 | 2026-08-24 | **Verified:** reports patterns over more than 200,000 keywords during 2025. **Proposed:** discuss longitudinal sampling and platform-change confounding. | D — proprietary sampling/parser and Google-specific scope. | B4; C5 | **Context-only** |
| V06 | [Profound Reports and Guides](https://www.tryprofound.com/reports-guides) | Vendor report hub | Current 2026 snapshot | 2026-08-24 | **Verified:** lists current AI-citation and visibility reports. **Proposed:** horizon-scanning queue; do not quote a report until its full methodology is audited. | D/E — hub metadata only; report-specific evidence ceilings unknown until review. | B1, B4; C5 | **Context-only** |
| V07 | [Adobe AI-Sourced Traffic Insights 2025](https://business.adobe.com/assets/pdfs/resources/reports/ai-traffic-report/ai-sourced-traffic-insights-2025.pdf) | Vendor analytics report / PDF | 2025 | 2026-08-24 | **Verified:** Adobe analytics report about AI-referred traffic and behavior. **Proposed:** contrast referral traffic with answer visibility, citation, and mention metrics. | D — customer/sample and attribution-method limits; traffic is not citation visibility. | B1, B4; C5 | **Context-only** |
| V08 | [One Year After Google AI Overviews Launched](https://videos.brightedge.com/assets/SGE-Guide/BrightEdge%20Report%20-%20AIO%20Overviews%20One%20Year%20Review%20Research%20Paper%20and%20Deep%20Dive%20.pdf) | Vendor research report / PDF | 2025-05 | 2026-08-24 | **Verified:** ten-page BrightEdge report based on its Generative Parser and dated one year after the AIO launch. **Proposed:** compare metric definitions and apparently conflicting vendor estimates; respect the report’s “not for distribution without consent” notice. | D — proprietary parser/sample, commercial framing, and no independent replication. | B4, B5; C5 | **Context-only** |
| V09 | [AI Search Visits Surging in 2025 — Organic Search Remains the Cornerstone](https://videos.brightedge.com/assets/blog/ai-search-visits-in-surging-2025/Industry%20Report%20Sep%202025.pdf) | Vendor industry report / PDF | 2025-09 | 2026-08-24 | **Verified:** eight-page report analyzing BrightEdge customer/query data from January–August 2025 and separating AI referrals from organic search. **Proposed:** traffic-attribution case study and a prompt to audit denominator, channel definition, and conversion window. | D — proprietary sample and attribution; referral traffic does not measure answer visibility or causal influence. | B1, B4; C5 | **Context-only** |

## 9. High-value videos and recorded lecture series

Videos are included for explanation and pacing. They should not carry unique scientific claims when a paper, standard, or official document is available.

| ID | English title / URL | Type | Date | Accessed | Verified fact → proposed use | Credibility / evidence ceiling | Supports | Status |
|---|---|---|---|---|---|---|---|---|
| M01 | [Stanford CS336: Language Modeling from Scratch — Spring 2025 Playlist](https://www.youtube.com/playlist?list=PLoROMvodv4rOY23Y0BoGoBGgQ1zmU_MT_) | Official university lecture playlist | 2025 | 2026-08-24 | **Verified:** recorded course sequence from foundations through training/evaluation topics. **Proposed:** prerequisite clips and lecture-pacing reference. | B — instructional recording; cite notes/papers for claims. | B2; C1 | **Candidate** |
| M02 | [Stanford CS25 Recordings](https://web.stanford.edu/class/cs25/recordings/) | Official university seminar recordings | Multi-year; current 2026 page | 2026-08-24 | **Verified:** includes “Retrieval Augmented Language Models” by Douwe Kiela among expert talks. **Proposed:** optional conceptual viewing with guided questions. | B — expert talks, not peer review. | B2, B6; C1, C3 | **Candidate** |
| M03 | [SIGIR 2024 Keynote: Representation Learning and Information Retrieval](https://www.youtube.com/watch?v=1ftmVDwdeO0) | Official conference keynote video | 2024-07-15 | 2026-08-24 | **Verified:** Yiming Yang surveys representation learning, dense IR, knowledge-enhanced retrieval, and RAG-related work. **Proposed:** field-history and research-frontier discussion. | B — expert synthesis; underlying papers support technical claims. | B2, B4; C1, C3 | **Candidate** |
| M04 | [Full Stack Deep Learning: Augmented Language Models](https://www.youtube.com/watch?v=YdeuQhlHmCA) | Practical course lecture video | 2023-05-11 | 2026-08-24 | **Verified:** explains retrieval, chaining, tools, and augmented-LM application structure. **Proposed:** systems overview before the implementation lab. | B/E — strong engineering pedagogy; examples and APIs are dated. | B2, B6, B8; C1, C3 | **Candidate** |
| M05 | [The GEO Community Live Kickoff](https://www.youtube.com/watch?v=7nrmVAbAQjY) | Practitioner community event recording | 2026-07-31 | 2026-08-24 | **Verified:** live practitioner/community event connected to The GEO Community. **Proposed:** capture contemporary terminology and practitioner questions, not technical evidence. | E — event discussion only. | B1; C1 | **Context-only** |
| M06 | [Stanford CS336: Lecture 1 — Overview and Tokenization](https://www.youtube.com/watch?v=SQ3fZ1sAqXI) | Official university lecture video | 2025 | 2026-08-24 | **Verified:** a clear example of opening a technically dense course with objectives, system scope, and implementation expectations. **Proposed:** production reference for the GEO course trailer/first lecture. | B — pedagogy and prerequisite framing only. | B8; C1 | **Candidate** |

## 10. Integration accounting and quality-control decisions

### Already absorbed into the project architecture

1. **CS336:** schedule-first density, visible assignments, prerequisites, and a build-oriented course rhythm.
2. **CS229 notes + arXiv:2506.02070 source audit:** coherent mathematical notation, explicit pedagogical roles, modular parts/appendices, and original rather than copied visual assets.
3. **Latest 42-paper CSV and normalized catalog:** canonical local literature corpus and module-routing layer.
4. **Original GEO paper and repository:** field origin, baseline visibility/intervention vocabulary, and a bounded reproduction target.
5. **CUHKGEO:** research-program framing across GEO, trusted AI, agents, and control/optimization; restrained deep-purple identity.
6. **FrontMind:** industry problem framing rewritten into the academically safer loop **observe → explain → intervene → validate**; no vendor performance claims are integrated.

### Highest-priority candidates for the next integration pass

1. **IR mechanism spine:** RAG, DPR, CS276/IIR, CMU Search Engines, and the SIGIR tutorial.
2. **Evaluation spine:** ALCE, TREC RAG, BEIR, Pyserini, and a claim-level citation rubric.
3. **Lab spine:** pinned Pyserini/BEIR retrieval lab; repeated-run visibility lab; controlled intervention/ablation lab; adversarial and governance clinic; reproducible evidence dossier.
4. **Operational boundary spine:** RFC 9309, sitemaps, first-party crawler docs, Bing AI Performance, and explicit separation of crawl → index → retrieve → cite → mention → refer.
5. **Governance spine:** NIST AI RMF/GenAI Profile, ISO/IEC 42001 metadata, current Chinese measures/standards, C2PA, privacy, and WCAG.

### Deliberately not promoted to scientific authority

- Vendor visibility percentages remain **Context-only** unless the exact sampling frame, parser, denominator, geography, date, and uncertainty are reproduced in the prose.
- Practitioner “best practices” are hypotheses or checklists until confirmed by primary research, first-party documentation, or a controlled experiment.
- `llms.txt` is a proposal, not a web standard; Google explicitly says it does not help or hurt Google Search visibility.
- Project 20252040-Z-469 and the Graph RAG project are not published standards as of 2026-08-24.
- A crawler request proves a request, not retrieval into an answer; a citation proves displayed attribution, not necessarily brand mention, rank, causality, or conversion.
- Structured data can improve machine-readable semantics and feature eligibility, but neither official Google guidance nor the cited quasi-experiment supports a universal AI-citation guarantee.

## 11. Refresh protocol before publication or course launch

1. Recheck every **C**, **D**, and **E** source within 30 days of release; record page update date, product surface, geography, account state, and screenshot hash when behavior matters.
2. Resolve every local paper to its current version, final venue if any, license, and figure reuse terms; retain the CSV/JSON snapshot as provenance.
3. Pin all repositories by commit and store environment lock files; archive only materials whose licenses permit it.
4. For every quantitative sentence, maintain a claim ledger with source, population, unit of analysis, denominator, uncertainty, time window, and evidence ceiling.
5. Label figures as **mechanism**, **observed data**, **reconstruction**, or **conceptual model**. Never draw a conceptual pipeline as though it discloses a closed platform’s internal architecture.
6. Run accessibility, link, citation, reproducibility, and legal/standards-status checks separately. Passing one gate does not imply passing another.

## 12. Legacy reconciliation and next catalog version

The legacy audit supplies three views that remain distinct:

1. **Current curated freeze (this file):** 86 controlled rows with status and
   evidence ceilings.
2. **Supplementary candidate queue:** first admit 15 additional primary papers
   and 16 official/platform/standards/governance sources; then audit 14
   reproducible repositories by commit, license, environment, and reference
   run. These are candidates for a future dated catalog, not hidden additions
   to the 86.
3. **Legacy watch/context corpus:** The GEO Community sitemap and 142 blog
   leads, 44 Bilibili links, three English videos, vendor/industry white papers,
   standardization projects, deleted or moved repositories, and duplicate
   manifestations. They support discovery, teaching, media design, claim audit,
   or status monitoring within their evidence ceilings; they do not become
   scientific authority through volume.

The first migration priority is primary research and official identity, not
more vendor prescriptions. Paper identities use DOI/arXiv/venue relations;
standards use issuer and identifier; repositories use owner/repository plus
commit; videos use stable media identity and timecode. Broken, recovered,
moved, and superseded states are appended as events rather than overwritten.

The ordinary-web-search continuation is recorded separately in
[Pass 2](reference_search_delta_2026-08-24_pass2.md),
[Pass 3](reference_search_delta_2026-08-24_pass3.md),
[Pass 4](reference_search_delta_2026-08-24_pass4.md), and
[Pass 5](reference_search_delta_2026-08-24_pass5.md),
[Pass 6](reference_search_delta_2026-08-24_pass6.md), and
[Pass 7](reference_search_delta_2026-08-24_pass7.md). Together they audit
additional repositories, tutorials, videos, curriculum comparators, standards,
regulatory and transparency sources, platform-version documentation, blogs,
white papers, paper-linked artifacts, and unresolved code/data gaps. None is
silently added to this 86-entry freeze.
