Local controlled anchors
Latest CSV · normalized catalog · mapping audit · TeX source audit
Provenance and project decisions only; never a surrogate for paper conclusions.
Controlled discovery universe
Domestic and international papers, courses, repositories, platform documents, protocols, standards, blogs, tutorials, reports, white papers, and videos—each separated into verified fact, proposed use, scope boundary, and project absorption status.
Source families
Latest CSV · normalized catalog · mapping audit · TeX source audit
Provenance and project decisions only; never a surrogate for paper conclusions.
GEO · RAG · DPR · ALCE · TREC RAG · CS276 · CS336
Mechanism, notation, evaluation, prerequisites, and curriculum models.
GEO-Bench · BEIR · Pyserini · RankLLM · MTEB · Ragas
Pinned implementations for bounded replication; code does not independently validate a result.
Google · Bing · Perplexity · RFC 9309 · Sitemaps · Schema.org
Named first-party behavior and protocol scope only; never a cross-engine guarantee.
NIST · ISO/IEC · C2PA · WCAG · PRC measures · GB/GB-T records
Status- and jurisdiction-aware controls; applicability and legal force remain explicit.
The GEO Community · Ahrefs · Semrush · Adobe · Stanford · SIGIR
Hypotheses, cases, pacing, and methods critique; primary sources carry scientific claims.
Complete ledger
Verified: 42 records, 87 fields, 42 DOI values, 41 URLs, 41 abstracts, and 42 local PDFs; SHA-256 c149a69140b0dcedc1855f8cbb02fc5cb1e1346470306c76cd68fbafc2e50468. Proposed: canonical reading-list input and claim-level paper queue.
Verified: machine-readable normalization of the latest CSV. Proposed: generate reference views, tags, module links, and version checks from one structured source.
Verified: preserves the initial 42-record triage and five metadata issues. The separate machine-readable sources/paper_route_crosswalk_v1.6.json is authoritative for final 28-chapter monograph, actual 12-chapter Core Notes citations, 16-week course, and evidence-card claim routes: 42 papers, 48 paper–claim edges, 17 unique claims, and three zero-edge cards. Proposed: retain the Markdown table as historical triage; use the crosswalk for v1.6 route checks.
Verified: tutorial-style notes with a compact title block, parts, appendices, equations, figures, and pedagogical environments. Proposed: emulate information architecture and teaching-role differentiation, not subject matter.
Open original source →Verified: 48 safe archive members; modular main.tex, seven parts, five appendices; SHA-256 b69a997d626dcbb94e3b91d83e3077251779f2d9ab767c4ea4209d84ea135fbf. The declared TeX Live 2025/pdflatex path reproduces an 84-page PDF; six page roles were visually sampled. The build remains warning-bearing, untagged, and has blank Title/Author metadata. License is CC BY-NC-ND 4.0. Proposed: transfer only abstract design principles via original prose and original figures, while enforcing stricter GEO release gates.
Verified: a schedule-first course page with lectures, assignments, readings, and recordings. Proposed: use its high-density weekly rhythm, assignment visibility, and prerequisite signaling.
Open original source →Verified: a 278-page coherent mathematical notes volume. Proposed: benchmark notation discipline, prerequisite refreshers, exercises, and appendix structure.
Open original source →Verified: the page presents research themes including GEO, agents/online decision-making, trusted AI, and control/optimization. Proposed: treat those claims as dated research-context questions and use restrained visual cues only.
Open original source →Verified: first-party industry framing organized around understanding, growth, and embedding. Proposed: retain the independently rewritten research loop “observe → explain → intervene → validate” and use brand cues sparingly.
Open original source →Verified: introduces GEO, GEO-Bench, visibility metrics, and controlled content interventions; reports gains up to 40% in its experimental setting. Proposed: field origin, baseline methods, and a reproduction/critique lab.
Open original source →Verified: official venue listing for the GEO paper. Proposed: use only to verify publication status and venue metadata.
Open original source →Verified: formalizes a parametric-plus-non-parametric retrieval/generation architecture and evaluates knowledge-intensive tasks. Proposed: canonical mechanism figure and notation for the retrieval-to-generation pipeline.
Open original source →Verified: evaluates dual-encoder dense retrieval for open-domain QA. Proposed: contrast sparse, dense, and hybrid candidate retrieval in the retrieval lab.
Open original source →Verified: evaluates correctness and completeness of citations for long-form generation. Proposed: distinguish answer quality, citation correctness, citation completeness, and source quality.
Open original source →Verified: long-running shared evaluation program; current tracks include retrieval-augmented generation. Proposed: model the course’s evaluation contracts, run files, qrels, and reproducibility package on TREC conventions.
Open original source →Verified: publishes topics, qrels, nuggets, and citation-support judgments. Proposed: adapt a small lawful subset or analogous original set for answer/citation scoring labs.
Open original source →Verified: documents sentence-level answer evaluation and citation-oriented judgments. Proposed: update the rubric beyond citation counts to claim-level support.
Open original source →Verified: covers indexing, retrieval, BM25, evaluation/NDCG, crawling, and web search. Proposed: prerequisite bridge and classical-IR comparison boxes.
Open original source →Verified: canonical treatment of indexing, scoring, evaluation, crawling, and link analysis. Proposed: prerequisite appendix and notation source.
Open original source →Verified: covers text search, RAG, and recommender/search systems in one contemporary course. Proposed: benchmark sequencing from classical retrieval to generative applications.
Open original source →Verified: combines LLM-based IR, vector search, reranking, conversational search, and query prediction. Proposed: cross-check advanced-course outcomes and workload.
Open original source →Verified: integrates IR/search with LLM-based conversational agents. Proposed: compare European learning outcomes and assessment balance.
Open original source →Verified: explicitly asks how to perform RAG without hallucination and retrieve from databases/knowledge graphs, with final projects. Proposed: inspiration for the course capstone and evidence-curation studio.
Open original source →Verified: tutorial materials and bibliography on generative retrieval. Proposed: scholarly bridge from document retrieval to generative IR; use slides only within their license.
Open original source →Verified: covers augmented language models, deployment, UX, and LLM operations. Proposed: derive production checklists and end-to-end project rhythm.
Open original source →Verified: demonstrates query reformulation, decomposition, retrieval, reranking, and multi-step validation. Proposed: optional agentic-retrieval lab with frozen dependencies.
Open original source →Verified: approximately two hours with videos and code on retrieval evaluation and advanced chunk/retrieval strategies. Proposed: use as optional preparatory practice, not as the course’s evidence authority.
Open original source →Verified: paper-linked code, benchmark assets, and evaluation implementation. Proposed: pin a commit, reproduce one bounded baseline, then document deviations and failures.
Open original source →Verified: heterogeneous IR evaluation across multiple datasets and retrieval systems. Proposed: small sparse-vs-dense-vs-hybrid lab and domain-shift discussion.
Open original source →Verified: reproducible sparse, dense, and hybrid first-stage retrieval over Lucene/Faiss with prebuilt indexes and qrels. Proposed: default retrieval lab toolkit with a pinned environment.
Open original source →Verified: supports reproducible pointwise/listwise LLM reranking workflows. Proposed: compare BM25, dense, and LLM reranking under fixed candidate sets and budgets.
Open original source →Verified: standardized evaluation across embedding tasks including retrieval. Proposed: teach model selection as dataset- and language-dependent rather than leaderboard universalism.
Open original source →Verified: implements dataset generation, RAG evaluation, and feedback workflows. Proposed: optional lab scaffold while separately validating metric definitions and judge reliability.
Open original source →Verified: modular LM programs and metric-driven compilation/optimization. Proposed: controlled multi-objective intervention lab with train/dev/test separation.
Open original source →Verified: supports test matrices, assertions, red-team checks, and CI workflows. Proposed: capstone regression suite for prompts, answers, citations, and policy checks.
Open original source →Verified: supports tracing and feedback functions for retrieval and generation. Proposed: optional observability comparison with Ragas; do not import vendor benchmark claims.
Open original source →Verified: a voluntary proposal for an /llms.txt file, not an IETF/W3C/ISO web standard. Google’s 2026 guide explicitly says Google Search ignores it. Proposed: use only in an experiment contrasting proposals with verified platform behavior.
Open original source →Verified: Google says existing SEO fundamentals apply; describes RAG and query fan-out; states indexing/serving are not guaranteed; says Google Search ignores llms.txt and no special AI markup is required. Proposed: primary platform-specific boundary text and myth-busting lab.
Open original source →Verified: mass generation without user value may violate scaled-content-abuse policy; emphasizes accuracy, quality, relevance, and appropriate creation context. Proposed: responsible publishing checklist.
Open original source →Verified: valid structured data creates eligibility for supported features but does not guarantee display. Proposed: schema validation lab with explicit “eligibility ≠ citation causality” language.
Open original source →Verified: preview reports citations, cited pages, sampled grounding queries, and trends across specified Microsoft surfaces; counts do not indicate rank, authority, or placement. Proposed: measurement-interface case study and metric-definition critique.
Open original source →Verified: recommends accurate XML sitemaps and lastmod; explicitly says no tool guarantees appearance in AI-generated results. Proposed: sitemap freshness exercise.
Open original source →Verified: marked content may remain indexed while being excluded from Bing snippets and AI summaries. Proposed: publisher-control lab contrasting crawl, index, retrieval, and display controls.
Open original source →Verified: distinguishes OAI-SearchBot access for search inclusion from GPTBot training controls and documents ChatGPT referral UTM parameters. Proposed: crawler-policy matrix and log-analysis lab.
Open original source →Verified: describes current web-search behavior, source review, query rewriting, and site eligibility; explicitly says placement is not guaranteed. Proposed: UI/surface observation protocol, with screenshots dated and labeled.
Open original source →Verified: distinguishes PerplexityBot and user-triggered retrieval behavior and publishes crawler information. Proposed: first-party source for crawler identification and access-log parsing.
Open original source →Verified: explains Perplexity’s stated treatment of robots controls. Proposed: compare declared policy with server-log observations without inferring intent from a user agent alone.
Open original source →Verified: normative syntax and matching rules for robots.txt. Proposed: authoritative protocol basis for all crawler-control examples.
Open original source →Verified: defines sitemap structure and limits. Proposed: technical discoverability lab and validation checklist.
Open original source →Verified: documents a shared structured-data vocabulary. Proposed: semantic entity/page representation exercise.
Open original source →Verified: specifies how sites notify participating engines of added, updated, or deleted URLs. Proposed: instrument notification latency and downstream crawl observations.
Open original source →Verified: normative accessibility success criteria. Proposed: mandatory quality gate for the course site, figures, forms, and labs.
Open original source →Verified: voluntary framework for incorporating trustworthiness into AI design, development, use, and evaluation. Proposed: structure governance around Govern, Map, Measure, and Manage while flagging the active revision.
Open original source →Verified: cross-sector companion profile for GenAI risks and actions. Proposed: risk taxonomy for evidence integrity, information security, content provenance, and incident response.
Open original source →Verified: specifies requirements for establishing, implementing, maintaining, and continually improving an AI management system. Proposed: organizational governance checklist and capstone audit trail.
Open original source →Verified: specifies cryptographically bound, tamper-evident provenance manifests; expressly does not judge whether provenance is “good” or “bad.” Proposed: multimodal provenance and disclosure lab.
Open original source →Verified: applies to specified public-facing generative AI services in mainland China and sets provider obligations. Proposed: jurisdiction-aware governance unit; obtain legal review for operational advice.
Open original source →Verified: requires specified explicit and implicit labels and related platform/provider duties. Proposed: publishing and provenance compliance checklist.
Open original source →Verified: current mandatory standard for AI-generated/synthetic content labeling methods. Proposed: standards crosswalk with S06 and C2PA, carefully separating mandatory labeling from optional provenance technologies.
Open original source →Verified: current recommended security standard for generative AI services. Proposed: security-governance matrix for a GEO measurement or optimization service.
Open original source →Verified: current recommended standard covering pre-training and fine-tuning data security. Proposed: data lineage, provenance, licensing, and contamination checklist.
Open original source →Verified: current personal-information security specification, with a revision process underway. Proposed: privacy controls for query logs, user studies, analytics, and model outputs.
Open original source →Verified: an approval-stage project, not a published national standard. Proposed: standards watchlist only; refresh status before publication.
Open original source →Verified: registered project concerning Graph RAG, not a published standard. Proposed: horizon note in the agentic/graph retrieval chapter.
Open original source →Verified: role-based map of papers, courses, tools, and implementation topics. Proposed: discovery index and competitor-curriculum checklist; verify every downstream claim at its original source.
Open original source →Verified: offers a conceptual contrast between link-oriented and synthesized-answer journeys. Proposed: introductory discussion prompt, paired with Google’s first-party guide and measurement evidence.
Open original source →Verified: practitioner-oriented explanation of A01. Proposed: optional pre-reading for nontechnical learners; all numerical claims cite A01 instead.
Open original source →Verified: provides configuration examples and crawler taxonomy. Proposed: debugging cases only after rechecking each user agent against vendor docs and RFC 9309.
Open original source →Verified: separates retrieval, generation, citation scoring, post-processing, and UI failure hypotheses. Proposed: convert into a falsifiable diagnostic worksheet anchored to ALCE, TREC RAG, and controlled logs.
Open original source →Verified: presents server-log parsing as an observation method. Proposed: privacy-safe lab distinguishing claimed bot identity, verified IP, request, response, and downstream outcome.
Open original source →Verified: proposes separating within-time repeated-sampling variance from between-time drift. Proposed: motivate repeated measures and time-blocked designs, with claims grounded in the local paper corpus.
Open original source →Verified: reports a very large Ahrefs-owned SERP sample and discloses important coverage limitations. Proposed: example of scale, sampling frames, and why headline percentages need denominators and dates.
Open original source →Verified: analyzes 863,000 SERPs and four million AI Overview URLs with an updated parser. Proposed: platform-drift case study and replication-design prompt.
Open original source →Verified: tracks 1,885 treated and 4,000 control pages and reports no major citation uplift across tested surfaces. Proposed: teach correlation vs treatment effects and challenge “schema guarantees citations.”
Open original source →Verified: distinguishes cited and mentioned outcomes using 3,981 domain appearances, 115 prompts, 14 countries, and four named surfaces. Proposed: metric-separation exercise, not a population estimate.
Open original source →Verified: reports patterns over more than 200,000 keywords during 2025. Proposed: discuss longitudinal sampling and platform-change confounding.
Open original source →Verified: lists current AI-citation and visibility reports. Proposed: horizon-scanning queue; do not quote a report until its full methodology is audited.
Open original source →Verified: Adobe analytics report about AI-referred traffic and behavior. Proposed: contrast referral traffic with answer visibility, citation, and mention metrics.
Open original source →Verified: ten-page BrightEdge report based on its Generative Parser and dated one year after the AIO launch. Proposed: compare metric definitions and apparently conflicting vendor estimates; respect the report’s “not for distribution without consent” notice.
Open original source →Verified: eight-page report analyzing BrightEdge customer/query data from January–August 2025 and separating AI referrals from organic search. Proposed: traffic-attribution case study and a prompt to audit denominator, channel definition, and conversion window.
Open original source →Verified: recorded course sequence from foundations through training/evaluation topics. Proposed: prerequisite clips and lecture-pacing reference.
Open original source →Verified: includes “Retrieval Augmented Language Models” by Douwe Kiela among expert talks. Proposed: optional conceptual viewing with guided questions.
Open original source →Verified: Yiming Yang surveys representation learning, dense IR, knowledge-enhanced retrieval, and RAG-related work. Proposed: field-history and research-frontier discussion.
Open original source →Verified: explains retrieval, chaining, tools, and augmented-LM application structure. Proposed: systems overview before the implementation lab.
Open original source →Verified: live practitioner/community event connected to The GEO Community. Proposed: capture contemporary terminology and practitioner questions, not technical evidence.
Open original source →Verified: a clear example of opening a technically dense course with objectives, system scope, and implementation expectations. Proposed: production reference for the GEO course trailer/first lecture.
Open original source →Reference record