Step 1 — Consolidated source pool (1c consolidator)

17 sources, S-1–S-17, from two blind searchers on disjoint data axes (epi-field: S-1–S-8; genomic-institutional: S-9–S-17). No duplicates found (checked; see note below). Grouped by line-of-evidence, each split by side — best-first within each side, ordering only, no scores (step 2’s job).

A. Case geography/timing & market-epicenter claim

B. Market wildlife trade & environmental/animal samples

C. Market-vs-lab ascertainment bias & data-suppression record

D. Official field investigation

E. Furin cleavage site

F. Closest known viral relatives (RaTG13/BANAL)

G. Lineage A/B & restriction-site molecular evolution

H. WIV documents & government/congressional investigations

Recurring datasets (expect shared data-basis D-nodes at step 2)

  1. Huanan market environmental-swab dataset — same raw China-CDC-collected swabs underlie S-2 (original) and S-3 (independent reanalysis); two independent interpretations of one dataset, not two datasets.
  2. Worobey 2022 case-geolocation compile (S-1) — a single assembled record (leaked/court-obtained/WHO-report fragments); later citations of its numbers (e.g. within S-7’s critique) aren’t independent new evidence.
  3. SARS-CoV-2 reference genome / GISAID sequence collection — underlies S-9, S-10, S-13, S-14 (each targets a different genomic feature/method, not a restatement).
  4. RaTG13 (S-11, WIV) vs BANAL (S-12, separate Laos expedition) — two distinct “closest known relative” datasets; do not conflate.

Exclusions (union of both searchers)

  1. Pekar, Worobey, Wertheim et al. 2022 Science (lineage A/B) — surfaced in epi-field slice, correctly deferred to genomic slice as S-13 (boundary call, not a duplicate: distinct DOI/method from S-1).
  2. Andersen et al. 2020 “Proximal Origin” — surfaced in epi-field slice, deferred to genomic slice as S-9.
  3. Furin-site/RaTG13/BANAL/DEFUSE/WIV-database material surfaced while mining the ACX post from the epi-field side — deferred to genomic slice, not individually itemized.
  4. WHO-China Joint Report — surfaced in genomic slice, correctly deferred to epi-field slice as S-6.
  5. Bloom 2021 (NCBI SRA deletion) — surfaced in genomic slice while distinguishing it from the WIV Sept-2019 takedown (S-17); correctly deferred to epi-field slice as S-5 (two different data-suppression events, two different institutions/times).
  6. Bloom, Virus Evolution 2023 — third independent reanalysis of the same Huanan swabs already covered by S-2/S-3; recurring-dataset cap, not noded; summarized in prose in S-3’s body.
  7. “A Critical Reexamination of Recovered SARS-CoV-2 Sequencing Data” (bioRxiv 2024/MBE 2025) — reanalysis of S-5’s same recovered sequences; not noded, summarized in S-5’s relevance_note.
  8. “Was Wuhan the early epicenter…? — A critique,” National Science Review — second, redundant critique of Worobey 2022; skipped to avoid triple-counting rebuttals (S-7 already covers this ground). Candidate for a second independent critique if step 2/3 wants one.
  9. US government intelligence assessments (ODNI Aug. 2021 summary + DOE “low confidence”/FBI “moderate confidence” ~2023) — named neutral/mixed anchor, content confirmed via search but not opened/minted; budget constraint. Flagged by its own searcher as the single highest-value addition if budget is revisited — the only candidate source giving a genuinely split verdict from one primary document rather than a one-sided argument.
  10. Alina Chan & Matt Ridley’s “Viral” (2021 book) — no genuinely original primary reporting found; treated as discovery hub only, not noded.
  11. DARPA’s DEFUSE rejection letter — genuine second primary document, folded into S-15’s motivatedness field rather than given its own node (confirms DEFUSE’s content/rejection, no new claim).
  12. The 3 debate videos + Weissman’s/Rootclaim’s full write-ups — not separately mined beyond named anchors already in the brief, time budget; a real coverage gap on both slices, not a considered rejection.

Note on near-collision (not a duplicate): S-1 and S-13 are both Worobey-coauthored Science 2022 papers with overlapping author lists and similar titles, but distinct DOIs, methods, and claims (case geolocation vs. lineage-divergence molecular clock) — correctly kept as two separate nodes per the search plan’s explicit boundary call, not merged.