Simon’s review of the COVID origins analysis
This review was written after the submission deadline.
Written by Claude (Claudian) recording Simon’s assessment plus a joint walkthrough of the analysis. Uses “Simon thinks…” for Simon’s own views; the concrete failure diagnoses come from a back-and-forth we did on the main report - Did SARS-CoV-2 first infect humans via zoonotic spillover or a lab leak.
Framing
- EpiStack is still a prototype, and there are probably plenty of flaws in the other analyses too — this review just happens to be about the one Simon dug into. COVID was predicted in advance to be the weakest case for the method, and it was.
- Simon’s broader stance: he mostly won’t try to make the epistemics work well for cases like COVID. The plan is to focus on questions where the evidence can be roughly trusted — not heavily motivated, and above all not fabricated — and ideally cases where there is little or no data, so the value added is the reasoning rather than the source-triage. COVID is the opposite: a small number of contested, motivated, plausibly-curated primary sources sitting inside an active cover-up.
Bottom line
- Simon thinks the analysis is quite bad — as a bottom-line probability it should not be trusted, and its ~83% zoonosis / ~17% research-related headline is the product of several independently zoonosis-favouring choices rather than a robust result.
- The one thing that does survive scrutiny is the conclusion that deliberate engineering is unlikely (the BANAL / natural-genome evidence is real, and it’s independent of the Chinese data and of any cover-up). Everything else — and especially the natural-spillover-vs-unmodified-collection-leak split that the top line hangs on — is soft.
Why this was expected (structural limits of the approach)
- No machinery for the social dimension. The pipeline is bottom-up: gather primary sources, weight them, do Bayesian updates. It has no built-in way to reason about the meta question — that evidence can be fabricated, that investigations can be fake, that a consensus can be manufactured. For COVID that’s not a detail; it’s most of the problem. A method that can only price the evidence in front of it will be systematically fooled by a motivated party that controls which evidence exists.
- Large N is not useful here, and probably hurt. Simon suspects that for a case like COVID — where there may only be a handful of primary sources that actually matter — pushing for a large evidence base (N≈20) mostly cluttered the analysis with unimportant, “both-sides”, conservatively-weighted material and diluted attention away from the few things that carry the answer. (He’s not certain of this, but it looks that way.) The effort profile is wrong: lots of small evidence nodes each nudged toward neutral, instead of a few hard Fermi estimates on the load-bearing cruxes.
- The trust baseline was lowered from 0.8 to 0.5 — deliberately. Simon had wondered whether the 0.8 threshold was kept. It was not:
agent-notes/curation.mdsetstrust_baseline = 0.50, with explicit reasoning that 0.8 “would drop the entire contested middle” (Pekar, Worobey, Crits-Christoph/Bloom, Andersen, Bruttel, the Bayesian analyses) and leave a biased high-trust-only set. That reasoning is coherent on its own terms, but it is exactly the wrong call for this kind of case: the whole point is that in a motivated/fabricated-evidence environment, letting in more dubious information is what screws up the model. Keeping 0.8 — and thereby admitting that most of the “evidence” here isn’t trustworthy enough to update hard on — would have been closer to right, even though it feels like it “answers a different question.”
What concretely went wrong (from the walkthrough)
f_wuhanwas done by relative anchoring, not a real Fermi. The single most load-bearing number in the whole analysis — the “why did it start in Wuhan?” factor — was set to a modest 5 by picking a point inside three authors’ relative multipliers (Rootclaim ~20×, Weissman ~80×, a “zoonosis-hub” ~2×). No clean absolute estimate of P(Wuhan | zoonosis) and P(Wuhan | lab leak) against city/reservoir baselines was ever computed. When we tried to do it properly, the model’s own numerator used a global set of labs (WIV ≈ 80% of worldwide sarbecovirus lab risk) while the chosenf_wuhanimplicitly used a China-only denominator — an inconsistent reference class. Fixing that (the zoonotic-emergence geography is southern China plus SE Asia, dense-city-detection-weighted) pushesf_wuhanto roughly 30–50, i.e. ~6–10× the value used. Since the research-bloc odds scale ~linearly withf_wuhanand the flip point is ~20, an honest value flips HC-1’s own posterior to research-majority (~65–70%). The headline is not robust to the one number the analysis itself calls the master crux.- Double-counting was invoked to keep
f_wuhanlow, which is having it both ways. HC-1 (the city-level “why Wuhan” factor, pro-lab) and HC-3 (the within-city “why the seafood market and not the WIV neighbourhood 12 km away” evidence, pro-zoonosis) were flagged as sharing a dependency, and the synthesis declined to multiply them. But those two operate at different spatial scales and point in opposite directions, so they’re closer to independent than “the same question”. Using the dependency as a reason to pre-deflatef_wuhanand still banking HC-3 as corroboration is double-dipping in the zoonosis direction. The clean handling is to pricef_wuhanhonestly (~40) and let HC-3 pull back on its own merits — then see where it lands (genuinely uncertain). - The Fermi estimates were bad in general.
f_wuhanis the worst offender, but the pattern recurs: load-bearing likelihood ratios (e.g. the DEFUSE furin-cleavage-site coincidence, priced at ~2.4:1) were set by anchoring and judgement rather than by decomposing the actual probability. When we decomposed the DEFUSE coincidence honestly — P(pandemic sarbecovirus has an FCS | zoonosis) ≈ 0.2, times P(a nearby aggressive bat-CoV proposal names FCS insertion | anything) ≈ 0.4 — the raw “spooky coincidence” comes out around 1-in-10, not the 1-in-thousands the lab-leak advocates claim, and also not something to wave away. The model happened to land in a defensible range there, but by luck of anchoring, not by construction. (Separately: the popular “P<0.002, only 1 of 800 sarbecoviruses has an FCS” and “1 in 4000” figures are the classic stacking of correlated weak factors and should be rejected — but the analysis rejecting them isn’t the same as the analysis doing the positive estimate well.) - Cover-up / fabricated-evidence was never priced, and it’s the crux. The analysis credits “investigations found nothing” as (weak) evidence, and treats the Chinese market and case-geolocation datasets as motivated-but-usable. It flags — but does not act on — the possibility that these were curated to manufacture a zoonosis signal, and that the absence of lab-side evidence is expected under an active cover-up rather than informative. If you take a bounded cover-up prior seriously, three things happen: (a) HC-3’s pro-zoonosis pull weakens (the very thing that would rescue zoonosis after an honest
f_wuhan); (b) the “absence of evidence” signals stop counting for zoonosis; (c) the genome/BANAL evidence is untouched (it’s an Institut Pasteur/Laos study, independent of China), so engineering stays ruled out regardless. The net is that a cover-up prior moves mass from natural spillover (H-41 - Natural zoonotic spillover with no research involvement) to the unmodified collection leak (H-42 - Research-related leak of an unmodified naturally-evolved virus (zoonotic collection)) — a wild virus WIV collected and leaked, then covered up — which is compatible with a natural genome, a market amplification, and the cover-up, and is therefore nearly evidence-proof. That scenario, not engineered bioweapon (H-43 - Research-related origin of a laboratory-manipulated virus (engineered lab leak)), is where the honest probability mass should be, and the bottom-up method has no way to get there.
How it could have gone right
- Simon’s sketch of the version that would have worked: keep the 0.8 trust threshold, so most of the motivated/contested material is admitted (if at all) as heavily discounted rather than used to update hard. Then have the model correctly recognise that there is cover-up activity, and therefore that the Chinese investigations and curated datasets can’t be trusted much. That leaves you with a small number of genuinely clear pieces of evidence — chiefly the public genome (BANAL rules out engineering; the FCS is the one anomaly) and the DEFUSE document — and at that point the whole game is doing a few actually-good Fermi estimates on those. That would probably have been reasonably good, and would have produced a wide, more lab-weighted distribution honestly centred on “if research-related, it’s a collection leak,” instead of a falsely precise 83%.
Process / rule notes
- Simon suspects the rule that a model is allowed to lower a trust score is possibly a bad rule (the pipeline both assigns per-source trust and lets downstream steps dock it further, e.g. CG-2 docked to 0.65, CG-11 to 0.80). Giving the model that discretion is a lever for motivated or just noisy adjustment, and it interacts badly with an already-too-low baseline.
- Relatedly, the definition of trust may need rewriting. Right now
trust_scoreis roughly “P(the key finding replicates cleanly), scored from method/statistics/corroboration only, explicitly not from believing the conclusion.” That “don’t look at the result” instruction is the right intent but is hard for a model to actually follow, and it’s not obviously the quantity you want anyway. Worth redesigning. - General direction: focus effort on the important parts of the evidence — the few cruxes that move the answer — rather than accumulating more peripheral, conservatively-weighted nodes. The current pipeline optimises for coverage and even-handedness; for hard cases that’s the wrong target.
Takeaways for the project
- Fermi estimation is a core competency that’s currently missing. Even Fable is, by default, still fairly bad at Fermi estimates and at getting the Bayesian bookkeeping right. This needs dedicated skills: how to do a Fermi estimate well, how to pick and hold a consistent reference class, how to decompose a factor further instead of anchoring on someone else’s multiplier, and how to avoid stacking correlated weak factors.
- There is a lot of headroom and prior art here — forecasting AIs exist and there may already be good skills/resources to adapt. Getting the Bayesian and Fermi layer right is important for every case, not just COVID, so it’s worth investing in even as we steer the project toward better-behaved (trustworthy-evidence, low-data) questions.