An independently authored, continuously-revised (through dozens of numbered versions) Bayesian analysis of SARS-CoV-2’s origin. Decomposes the question into likelihood ratios for the outbreak’s Wuhan location, the sarbecovirus lineage, the furin cleavage site, and the DEFUSE grant proposal, then combines them; the versions found convert to combined odds ranging roughly 4:1 to 12:1 (and in earlier/other framings up to ~165:1) favoring lab leak, i.e. more confident in lab leak than the debate’s judges were in zoonosis, but explicitly less extreme than Rootclaim’s own number. Scott Alexander’s ACX piece notes he did not include Weissman in his six-way comparison table “because it would have taken too long to translate his language into mine” — so Weissman’s analysis is an important 7th/parallel data point outside the “six,” not one of them.

relevance_note: named directly in the case brief as an anchor Bayesian analysis; one of the most technically detailed lab-leak-favoring analyses outside Rootclaim itself, and the direct target of judge Eric’s own follow-up rebuttal.

Extracted (structured summary)

Bayesian-synthesis artifact resting on no data of its own (data_basis: []); its force lives in the hypotheses (overall conclusion, prior, ascertainment claim) and the factor/combination arguments that feed the step-7 prior. Numbers are version-dependent (continuously-revised Substack).

Overall conclusion

H-17 - SARS-CoV-2 most likely originated from a research-related incident rather than natural zoonosis (Weissman synthesis)

Top-line conclusion of the artifact; specific odds vary by version (continuously revised). Rests on no data of its own - its force comes from the factor arguments and prior below.

Link to original

Prior / base rate

H-18 - The prior probability of a research-related pandemic origin in 2019 is on the order of 1 in 100 to 1 in 200

Link to original

A-37 - Historical lab-leak base rates plus WIV-specific factors set the 1-in-100 to 1-in-200 prior

Reasoning

Literature estimates put the probability of a major human-transmissible leak at roughly 0.2-1% per year per relevant BSL-3 laboratory. WIV-specific factors raise concern: SARS-related coronavirus work was conducted partly at BSL-2 (inadequate for gain-of-function on such viruses), and State-Department cables had warned of safety deficiencies. Conditional on DEFUSE-type work actually proceeding, Weissman puts P0if(2019,LL) ~= 1/100. Multiplying by an estimated ~50% probability that the work actually occurred (DEFUSE was formally rejected by DARPA, but subsequent Chinese-Academy funding and Shi Zhengli’s non-denial keep the probability substantial) gives P0(2019,LL) ~= 1/200, i.e. starting prior odds around 1/70 favouring zoonosis. He flags these as ‘obviously very rough’, possibly off by a factor of ~10. This base-rate estimate is the L0 term feeding the combination in H-17.

Validity

Reconstruction. Premises: per-lab-per-year leak base rate ≈ 0.2–1%; WIV concentrated the relevant SARSr-CoV work (partly at BSL-2, with warned safety deficiencies); conditional on DEFUSE-type work proceeding, P0if(2019,LL) ≈ 1/100; and ≈50% probability the work actually occurred. Load-bearing step: a law-of-total-probability composition, P0(2019,LL) ≈ P(LL | work)·P(work) ≈ (1/100)·(1/2) ≈ 1/200, giving a prior in the 1/100–1/200 band.

Checked. The arithmetic is elementary and correct (1/100 × 0.5 = 1/200; the 1/100–1/200 range spans the work-conditional and work-marginal figures). The composition is a valid total-probability step and is if anything conservative, since it drops the (small) no-DEFUSE leak path rather than inflating the estimate. No completion-by-force: the inputs are granted premises, and no undercutting defeater breaks the link from these inputs to an order-1/100–1/200 prior. Whether the base rate and the 50% figure are true is a premise-truth matter priced at step 7 (and Weissman himself flags them “obviously very rough,” possibly off by ~10×); that does not touch the validity of the inference. Approved.

Link to original

Likelihood factors

A-38 - The outbreak occurring specifically in Wuhan is a strong likelihood factor for lab origin

Reasoning

Under a zoonotic-wildlife hypothesis, a spillover could occur anywhere along China’s (and SE Asia’s) wildlife-trade geography, of which Wuhan handles a tiny fraction; Wuhan’s markets handled under 0.01% of China’s mammalian wildlife trade (raccoon dogs perhaps 1/2000 of national trade). Conditioning on the outbreak starting in Wuhan therefore removes ~99% of the zoonotic probability mass. Under a research-related origin, the relevant work (WIV and Wuhan CDC) was concentrated in Wuhan, so a Wuhan onset is near-expected. The ratio of these conditional probabilities gives a large positive logit, L2 ~= 4.4, in favour of lab origin. This is one of the dominant terms in the H-17 synthesis; it is partially dependent on the market-ascertainment issue (see the ascertainment argument), which is why it is not multiplied naively with the case-location evidence.

Original

statement: “That the pandemic began in Wuhan - a city with under 1% of China’s population but home to the main labs studying SARS-related bat coronaviruses - is far more expected under a lab origin than under zoonosis, contributing a large positive likelihood factor (about 4.4 logits).”

Validity

Reconstruction. Premises: WIV/Wuhan-CDC work was concentrated in Wuhan (→ P(Wuhan onset | lab) near 1); Wuhan handles a tiny share (<0.01%) of China’s mammalian wildlife trade. Load-bearing step: because zoonotic spillover “could occur anywhere along the wildlife-trade geography,” conditioning on Wuhan removes ~99% of zoonotic probability mass, so the ratio P(Wuhan|lab)/P(Wuhan|zoonosis) ≈ e^4.4 ≈ 81.

Corrected. The step from “Wuhan’s trade share is tiny” to “conditioning on Wuhan removes ~99% of zoonotic mass” requires the hidden premise that zoonotic-spillover probability is proportional to raw wildlife-trade volume. An undercutting defeater survives without denying the trade-share premise: spillovers that seed a detected pandemic are weighted toward large, dense, well-connected cities with early detection capacity (Wuhan is an 11M-person hub whose markets sold live susceptible mammals), so P(zoonotic outbreak first detected in Wuhan) can be far higher than Wuhan’s raw trade share — this is precisely the contested crux in the Worobey/Pekar-vs-Weissman dispute. The defeater undercuts the magnitude (4.4 logits) but not the direction: the labs are pinned to Wuhan while a zoonotic origin retains geographic freedom, so P(Wuhan|lab) > P(Wuhan|zoonosis) still holds and the factor still favours lab. Per the step-06 pattern (a strong-magnitude / “far more expected” claim that survives only as an “is evidence for”), corrected to the direction-plus-caveated-magnitude form. Checked — I traced the reference-class dependence myself; it does not rest on author or venue.

Link to original

A-39 - The sarbecovirus lineage, furin cleavage site, and DEFUSE-matching features each favour lab origin

Reasoning

Three virus-feature factors each condition on the specific pre-2020 research plan. (a) Sarbecovirus type: only about one of ~19 recent emerging pathogens in China was a sarbecovirus, whereas the DEFUSE/WIV program pre-specified exactly this virus category, giving L1 ~= 2.3 in favour of lab origin. (b) Furin cleavage site: SARS-CoV-2’s FCS is rare or absent among natural sarbecoviruses, and its nucleotide-level features resemble patterns used in engineered constructs, matching the FCS-insertion plan written into the DEFUSE proposal - contributing roughly 2-3 logits, though with wide uncertainty. (c) Restriction-enzyme pattern: the genome’s restriction-site spacing matches patterns in DEFUSE drafts and is uncommon in natural viruses, contributing roughly 1-2 logits, also uncertain. Because each is estimated conditional on the same DEFUSE plan, their uncertainties are large and they are discounted toward 1 in the combination (per the combination argument), but their central estimates all point the same way, reinforcing H-17.

Validity

Reconstruction. Premises (assumed): (a) sarbecoviruses are ~1/19 of recent emerging pathogens in China, while the DEFUSE/WIV program pre-specified this category; (b) the FCS is rare/absent among natural sarbecoviruses and its nucleotide features resemble engineered constructs, matching DEFUSE’s FCS-insertion plan; (c) the restriction-site spacing matches DEFUSE drafts and is uncommon in natural genomes. Load-bearing step: for each feature, P(feature | research plan) > P(feature | zoonosis), so each is a separate likelihood factor favouring lab.

Checked. Conditional on each stated premise, the direction is a valid likelihood inference — a feature that is common under the pre-specified research plan but rare in the natural comparison class carries LR > 1 for lab origin. The word “independently” in the statement is read charitably as “each on its own points toward lab” (not statistical independence); the body itself notes the three share DEFUSE-conditioning and are discounted toward 1 in combination, so no double-counting is asserted here — that guard lives in A-36. Probed defeaters (the RE-site pattern is not actually statistically anomalous once natural genomes are surveyed; the FCS’s “engineering-like” character is disputed) attack the truth of premises (b)/(c), not the inferential step, and are priced at step 8. No premise-preserving defeater breaks the direction of any of the three factors. Approved.

Link to original

Market ascertainment

H-19 - The Huanan-market early-case clustering is largely an ascertainment-bias artifact, not evidence of a market spillover

Link to original

A-40 - Ascertainment bias, not a market spillover, best explains the Huanan-market case clustering

Reasoning

Early case-finding disproportionately tested people with a market link, so market-associated positives were over-represented, creating the illusion that the market was the source (everyone testing positive had been there). Weissman marshals additional strands against a market spillover: (i) the more-ancestral lineage A appeared in locations away from the market while the derived lineage B clustered at the market - the reverse of the ordering expected if the market were the spillover site; (ii) SARS-CoV-2 RNA in market environmental swabs showed no tendency to co-occur with DNA of candidate non-human host species, unlike genuine animal coronaviruses whose RNA correlated with their hosts’ DNA (Bloom), undercutting an infected-animal source at the market; (iii) early social-media (Weibo) illness reports clustered on the south side of the Yangtze near the Wuhan CDC and WIV rather than the market; and (iv) Levin’s spatiotemporal Bayesian reanalysis found ~27:1 odds favouring a research-side source over the market. Together these support H-19 (the clustering is an ascertainment artifact) and, by removing the market-zoonosis evidence, shift the overall odds in H-17 toward a research-related origin.

Validity

Reconstruction. Premises (assumed): early case definitions preferentially tested people with a market link; (i) ancestral lineage A appeared away from the market while derived lineage B clustered at it; (ii) market environmental RNA did not co-occur with candidate host DNA (Bloom); (iii) early Weibo illness reports clustered near CDC/WIV; (iv) Levin’s spatiotemporal reanalysis gave ~27:1 for a research-side source. Load-bearing step: market-link-biased testing manufactures a market-centred cluster, and (i)-(iv) each remove or reverse the market-spillover reading, so ascertainment bias is the better explanation and, with market-zoonosis evidence subtracted, the odds shift toward lab.

Checked. The ascertainment mechanism validly produces over-representation of market-linked positives conditional on the premise that early testing was market-biased. Each strand is a valid conditional inference in the stated direction: (i) an ancestral-lineage-away / derived-at-market pattern is the reverse of the market-as-source expectation; (ii) RNA-without-host-DNA undercuts an infected-animal source; (iii)/(iv) directionally support a non-market/research-side locus. The comparative conclusion (“better explained by ascertainment than by true spillover”) follows once these strands neutralise the market evidence. Probed the strongest defeater — Worobey/Pekar’s finding that early cases without market links also cluster spatially on the market, which market-link-biased testing alone would not generate, and Pekar’s two-introduction reading of the lineage A/B split. These contest the truth of the ascertainment premise and of strand (i), not the validity of the inference conditional on the stated premises, so they are priced at step 8; strand (iii) (Weibo/CDC proximity) is individually weak but only additive here. No premise-preserving defeater blocks the conditional step. Approved.

Link to original

Combination

A-36 - Logit-additive combination with dependence handling and uncertainty discounting yields combined odds favouring lab leak

Reasoning

Weissman writes the posterior log-odds as a sum of logit contributions, ln(P(LL)/P(ZW)) = L0 + L1 + L2 + …, with approximate values: prior L0 ~= -4.2 +/- 2.3 (about 1/70 favouring zoonosis), sarbecovirus-type L1 ~= 2.3, Wuhan-location L2 ~= 4.4, furin cleavage site ~= 2-3, and a restriction-enzyme / DEFUSE-matching pattern ~= 1-2. Dependence between factors is handled explicitly: the case-location data and the phylogenetic data are not treated as independent updates when they bear on the same particular version of the market-zoonosis hypothesis (the location data supported a version the phylogenetic data make implausible), so they are not multiplied as if independent. Uncertainty is propagated as distributions on each logit: uncertainty in a likelihood ratio discounts that ratio toward 1 but does not discount the prior, whereas uncertainty in the prior discounts both - so integrating over realistic uncertainty pulls extreme point odds back toward 1. The net still leaves lab leak favoured: point estimates drive P(zoonosis) below ~1%, and the uncertainty-integrated odds remain single-digit to low-double-digit favouring a research-related origin. This is the combination machinery that produces H-17’s headline odds.

Validity

Reconstruction. Premises: the individual logit values (L0 ≈ −4.2, L1 ≈ 2.3, L2 ≈ 4.4, FCS ≈ 2–3, RE ≈ 1–2); that dependent factors are not multiplied as independent; and that uncertainty is propagated as distributions on each logit. Load-bearing step: Bayes in odds form is additive in log-odds (ln posterior-odds = ln prior-odds + Σ ln LRᵢ for conditionally-independent evidence), so combining these logits and integrating over uncertainty yields net odds whose sign favours lab.

Checked. (1) The additivity itself is a valid theorem of Bayesian updating; the dependence caveat (not multiplying the case-location and phylogenetic factors, which bear on the same market-zoonosis sub-hypothesis) is the correct guard against double-counting and moves in the right direction. (2) The uncertainty claim — “uncertainty in a likelihood ratio pulls that ratio toward 1 but not the prior; uncertainty in the prior pulls both” — is defensible read as marginalizing the posterior probability over parameter uncertainty: E[P(zoonosis)] of a bounded quantity is pulled toward less-extreme values, i.e. odds toward 1. (3) The conclusion’s sign is robust conditional on the premises: the positive logits (≈2.3+4.4+2.5+1.5 = 10.7) dominate the prior (−4.2) by a wide margin, so any symmetric discounting toward 1 leaves the net favouring lab without flipping.

Probed defeater: L1/FCS/RE all condition on the same DEFUSE plan, a shared dependence that could inflate the sum beyond what the explicitly-handled case-location/phylogenetic dependence covers. This bears on the magnitude of the net odds, not the validity of the additive-combination step or the sign of the result (priced at step 7/8), so it does not undercut the stated conclusion. Approved.

Link to original