# 普通网页搜索增量 · Pass 6

**核验日期：** 2026-08-24（Asia/Shanghai）  
**方法：** 普通网页搜索与一手页面逐项核验；未使用 research-assistant 工作流  
**冻结边界：** 本轮不静默改变 GEO-42 论文分母或 86-source Reference Universe；先进入 post-freeze delta，经身份、版本、全文、许可和课程/专著编辑审计后再正式 commit。

## 1. 本轮结论

本轮与 86-source freeze、Pass 2–5、post-freeze queue 和最新 42-paper
catalog 逐题名比对，保留 25 个无精确重复的候选：16 项论文/预印本/工作
论文，7 项官方标准、平台、GitHub 或技术材料，2 项大学课程。优先吸收的
不是新的“GEO 技巧”，而是四个方法缺口：引用与归因统一评价、冲突证据、
可复现报告标准、以及从检索/引用到推荐与行为结果的分层审计。

状态词含义：

- **Design-absorbed:** 已纳入目录、Lab 或质量门的设计思路，但尚未进入冻结 References 分母。
- **Candidate:** 值得全文/代码审计后纳入。
- **Context-only:** 可用于案例、反证或行业状态，不承担中心机制结论。

## 2. 正式论文与研究预印本

| ID | 来源 | 已核验增量 | 拟吸收位置 | 证据上限 / 当前状态 |
|---|---|---|---|---|
| P6-01 | [Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models](https://aclanthology.org/2026.acl-long.1430/) · Schreieder, Schopf & Färber · ACL 2026；[数据仓库](https://github.com/faerber-lab/AttributeCiteQuote) | 134 篇论文、300 个指标、七维 taxonomy；将 attribution / citation / quotation 放入同一评价地图 | 专著 Ch5–6 的指标地图；课程 W05/W07；将仓库 CSV 设计成可审计 reading map | ACL 正式论文；与 ALCE/TREC RAG 近邻但不重复。**Design-absorbed / priority commit.** |
| P6-02 | [CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation](https://aclanthology.org/2026.acl-long.282/) · Choi et al. · ACL 2026 | 区分错误引用、漏引和替代合法引用；评价人类引用选择之外的 valid alternative | W05 claim–citation Lab 扩展；加入 abstention 与 alternative-valid-citation 标签 | 数据集/领域有边界；不作为通用引用正确率。**Candidate.** |
| P6-03 | [ReproEvalCard](https://aclanthology.org/2026.acl-short.22/) · Pattnayak & Bhatia · ACL 2026 | 审计 55 篇 LLM pipeline 论文，提出 prompt、judge、corpus snapshot、tool schema、trace、randomness 等报告项 | Capstone evidence dossier；Source Ledger；所有实验的 minimum reporting contract | 报告标准，不是 ISO/NIST 标准。**Design-absorbed / priority commit.** |
| P6-04 | [Redefining Retrieval Evaluation in the Era of LLMs](https://aclanthology.org/2026.eacl-long.391/) · Trappolini et al. · EACL 2026 | UDCG 同时表示对生成有用与可能干扰的检索结果 | Ch4/W04 metric card；与 nDCG 进行“retriever consumer changed”对照 | 五数据集/六模型结果不能直接成为 GEO 总分。**Design-absorbed / priority commit.** |
| P6-05 | [ConfRAG: Benchmarking LLM Reasoning over Conflicting Web References](https://aclanthology.org/2026.acl-long.11/) · Yuan et al. · ACL 2026；[代码/数据](https://github.com/XaiverYuan/ConfRAG) | 1,814 个真实问题、平均 9.58 段证据、57.2% 含显式冲突 | W05/W10 冲突保留 Lab；理由覆盖、冲突检测和 refusal gate | 重建网页数据会漂移；固定 snapshot 后使用。**Design-absorbed / priority commit.** |
| P6-06 | [Context Attribution with Multi-Armed Bandit Optimization](https://aclanthology.org/2026.findings-acl.565/) · Pan et al. · Findings ACL 2026；[CAMAB](https://github.com/pd90506/camab) | 组合 bandit 估计段落影响并降低调用量 | 与 MaxShapley 做效率/假设/预算对照；W07 extension | API/log-probability 依赖；不替代因果 attribution。**Candidate.** |
| P6-07 | [Media Source Matters More Than Content](https://aclanthology.org/2025.emnlp-main.872/) · Dai et al. · EMNLP 2025 | 控制内容、改变媒体来源标签，分离 content-position 与 source-identity preference | Ch6/W14 偏差案例；来源标签作为实验处理变量 | 限政治新闻数据与所测模型。**Design-absorbed / priority commit.** |
| P6-08 | [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483) · Zhang et al. · 2025-12 | 55,936 查询、六个 LLM 搜索引擎与 Google/Bing；覆盖、可信度、政治中立、安全与页面特征 | W07/W10 跨系统 observation panel；vendor 数字的独立对照 | 预印本、动态平台观察；不推断长期或因果效果。**Design-absorbed / priority full-text audit.** |
| P6-09 | [Auditing Citation Behavior in AI-Generated Search Summaries](https://proceedings.mlr.press/v318/kakimov26a.html) · Kakimov et al. · PMLR 2026 | 按 retrieval rank 与 provenance 条件化引用审计；Google AI Overviews YMYL case | Ch5–7 条件概率图；W07 stage-resolved audit | 单一平台和时间窗。**Design-absorbed / priority commit.** |
| P6-10 | [The Attribution Crisis in LLM Search Results](https://www.cambridge.org/core/journals/data-and-policy/article/attribution-crisis-in-llm-search-results-estimating-ecosystem-exploitation/170DD0B88E5F5AEA8F69F2E9AF1328E3) · Strauss et al. · Data & Policy 2026；[复现仓库](https://github.com/AI-Disclosures-Project/Ecosystem_Exploitation_In_Search_Results) | 约 13,929 条 Search Arena 记录；定义 retrieved-but-uncited attribution gap；开放清洗数据与代码 | Ch6/W07 telemetry 与 disclosure case；区分 provider log 与完整候选集 | provider log 可能预过滤；“exploitation”需按作者定义限缩。**Design-absorbed / priority commit.** |
| P6-11 | [Synthetic Sources?](https://arxiv.org/abs/2605.23684) · Allaham & Diakopoulos · 2026-05 | 712 个真实查询、四个生成式搜索系统、三类高影响主题；研究来源集中与疑似 AI 内容 | W10 source-ecosystem audit；检测器不确定性案例 | “AI-generated”依赖检测器，不是真实作者身份 ground truth。**Candidate.** |
| P6-12 | [Auditing Preferences for Brands and Cultures in LLMs](https://arxiv.org/abs/2603.18300) · Rienecker et al. · 2026-03 | ChoiceEval 把自由回答规范为 top-k choice set，并引入 persona / brand / culture 层 | W06 query frame、W14 fairness、recommendation audit | 预印本；国家/文化归属解释需审查。**Candidate.** |
| P6-13 | [Whose hotel does the AI recommend?](https://arxiv.org/abs/2606.16344) · Baig, Gillani & Ali · 2026-06 | 预注册 choice-based conjoint；独立随机价格、评分、评论、认证与位置；估计推荐概率 AMCE | W08 因果模板；推荐阶段可复现 synthetic Lab | 合成酒店与模型快照限制外推。**Candidate.** |
| P6-14 | [Whose doctor does the AI recommend?](https://arxiv.org/abs/2608.14399) · Gillani & Baig · 2026-08 | 3,024 choice sets、七模型、40,068 响应；声誉、人口属性和首位效应 | W14 高风险推荐、公平性与 position-effect case | 预印本；不作医疗建议依据；可与 P6-13 合并阅读。**Candidate.** |
| P6-15 | [The Price of Advice](https://www.econstor.eu/bitstream/10419/336742/1/194960134X.pdf) · Zac & Gal · Working Paper 375 · 2025-11 | 随机分配传统搜索/GPT/Gemini/定制 GPT，观察购买与支出 | Ch1 visibility→behavior outcome ladder；W08 downstream outcome | 工作论文；不写成普遍消费效应。**Candidate.** |
| P6-16 | [ChatGPT as a News Recommender System](https://epub.ub.uni-muenchen.de/135475/) · Schatto-Eckrodt et al. · SocArXiv v3 · 2025-11 | 直接比较 Web 与 API 的来源类型、多样性和许可媒体出现 | W06/W07 将 interface/access path 作为处理变量 | 预印本和时点依赖。**Candidate.** |

## 3. 官方标准、平台、GitHub 与技术材料

| ID | 来源 | 已核验增量 | 拟吸收位置 | 证据上限 / 当前状态 |
|---|---|---|---|---|
| P6-17 | [ISO/IEC TS 42119-2:2025 — Testing of AI, Part 2](https://www.iso.org/standard/84127.html) | AI 测试生命周期、风险、过程、文档、stakeholder、test level/type/design | 全书/全课 QA taxonomy；外审矩阵 | 仅公开 metadata/abstract/目录；正文付费且未取得。**Design-absorbed / official-metadata only.** |
| P6-18 | [NIST AI 200-2 Initial Public Draft — TEVV-Athlon](https://www.nist.gov/artificial-intelligence/ai-research/tevv-athlon-framework-evaluating-ai-systems) · 2026-08-07 | objective → event/tool → block/measure → evidence 的四阶段定制 TEVV | Capstone、实验计划、evidence dossier 和 stop gate | 必须标作 initial public draft。**Design-absorbed / priority commit.** |
| P6-19 | [Powering Product Discovery in ChatGPT](https://openai.com/index/powering-product-discovery-in-chatgpt/) · OpenAI · 2026-03-24 | 商品发现/比较界面和 ACP 商品目录接入的一手边界 | Ch3/W10 product discovery surface；官方界面案例 | 不证明排序公式、收录保证或效果。**Design-absorbed / official context.** |
| P6-20 | [Agentic Commerce Protocol](https://github.com/agentic-commerce-protocol/agentic-commerce-protocol) · OpenAI/Stripe | Apache-2.0；版本化 OpenAPI、JSON Schema、示例、changelog、governance | W10 product-feed/schema validation GitHub Lab | beta/open protocol，不是 ISO/IETF 标准。**Design-absorbed / priority repository audit.** |
| P6-21 | [Claude Platform — Web Search Tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool) · Anthropic | citations、domain controls、dynamic filtering、response inclusion 和显式 tool version | W03/W07 platform-version experiment；补 Anthropic 一手资料 | 只支持公开 API 行为。**Design-absorbed / official platform record.** |
| P6-22 | [The Crawl-to-Click Gap](https://blog.cloudflare.com/crawlers-click-ai-bots-training/) · Cloudflare Radar · 2025-08-29 | 固定 cohort 区分 training/search/user-action crawler，比较 crawl 与 referral | Ch3/W03/W07 stage vocabulary 与 denominator audit | Cloudflare 客户专有观察数据；不作因果结论。**Design-absorbed / context with telemetry ceiling.** |
| P6-23 | [How Repeatable Are AI Recommendations?](https://llmaudit.app/research/how-repeatable-are-ai-recommendations) · Soria · 2026-08-23 | 30 类目×10 问法×3 provider×5 次；公开 CC BY 4.0 cell-level data | W07 recommendation stability mini-Lab | 单日、预算模型、温度不一致、8.6% 抽取失败、抽取代码私有。**Context-only / reproducible-data candidate.** |

## 4. 大学课程、讲义与公开视频

| ID | 来源 | 已核验增量 | 拟吸收位置 | 证据上限 / 当前状态 |
|---|---|---|---|---|
| P6-24 | [Stanford CS124 — From Languages to Information](https://web.stanford.edu/class/cs124/) · Dan Jurafsky · Winter 2026 | 公开 IR/RAG 章节、IR Lab、GitHub starter/solutions、embeddings、agents 与 recommender | W03–W04 先修桥；“讲义—Lab—代码—答案”制作基准 | 教学参照，不是 GEO 效果证据。**Design-absorbed / priority curriculum comparator.** |
| P6-25 | [UMD CMSC 848Q — Good AI Answers to Questions](https://www.cs.umd.edu/~jbg/teaching/CMSC_848/) · Jordan Boyd-Graber · Spring 2026 | 公开 YouTube、slides、GitHub 作业；IR、dense retrieval、generation、fact checking、冲突证据、评价陷阱 | W03–W08 公开视频/课程编排；“答案为何会错”叙事参考 | 只有公开材料进入资源页；登录受限 Panopto 不列作开放工件。**Design-absorbed / priority curriculum comparator.** |

## 5. 正式纳入前的提交门

1. 冻结论文/标准/仓库的版本、DOI/venue、发布日期和合法获取路线。
2. 对 P6-01–P6-10 做全文 evidence card；数字必须绑定数据集、系统、时间窗和不确定性。
3. 对 GitHub 仓库记录 commit、license、依赖、最小 smoke test 和网络/数据下载边界。
4. 将 ISO/NIST draft/TS/standard 身份分开；未取得的付费正文不得被推断。
5. 视频/课程材料只承担教学组织、解释和观看指引；技术结论回链到论文/官方文档。
6. 下一次 Reference Universe commit 必须显式更新分母、状态、课程/专著位置、source ledger 和 public release manifest；本文件本身不改变 86-source freeze。
