Identifies what Weissman describes as a basic error in the Bayes-factor calculation in Pekar et al. 2022 (S-1): the reported comparison conflates likelihoods computed under non-equivalent conditionalizations for the one- vs. two-introduction models. He shows that correcting this reverses the direction of the conclusion — the corrected calculation favors a single introduction over two. Notes that even the 2024 Science-issued erratum to S-1 (which fixed a separate coding bug) left this deeper framing error unaddressed, so the version of the paper on Science’s website continues to report Bayes factors Weissman argues are wrong in kind, not just in degree. Converges independently with McCowan (S-12) on “the published two-introduction Bayes factor doesn’t survive a correct calculation,” via a different diagnosis of the error.

relevance_note: the most sustained and technically detailed rebuttal of Pekar et al.’s central quantitative result, from an author with a documented stake in the debate’s outcome.

Extracted (structured summary)

Reported result (rests on S-1)

O-21 - Pekar 2022 reported two-introduction Bayes factor fell from ~60 to ~4.3 after the 2024 coding-error erratum

Weissman re-reads Pekar et al. 2022 (P2022) as the data basis; these figures describe P2022’s own reported result (original ~60, post-erratum ~4.3), not a new dataset.

Link to original

The conditionalization error (rests on S-1)

O-22 - Pekar 2022 computed the two-introduction likelihood by squaring only the polytomy criterion, omitting the size-ratio and sequence-difference conditions imposed on one introduction

Records the structure of P2022’s calculation (from P2022’s own simulations and supplement): I2’s likelihood used only the polytomy sub-criterion tauP applied twice, while I1’s likelihood additionally required the size-ratio and sequence-difference conditions. This asymmetry is the factual basis for Weissman’s error argument.

Link to original

Correction and reversal

A-35 - Correcting Pekar 2022 asymmetric conditionalization reverses its Bayes factor to at least 4.4 favouring a single introduction

Original

Requiring the two-introduction model I2 to satisfy the same size-ratio and sequence-difference conditions that Pekar et al. 2022 imposed on the one-introduction model I1 collapses P(tau|I2) from 0.226 to at most 0.0071, turning the reported ~4.3 Bayes factor favouring two introductions into a factor of at least 4.4 favouring a single introduction.

Validity assessment — corrected (checked)

Reconstruction. Load-bearing step: the same observed data properties must condition both hypotheses; Pekar conditioned I1 on the full property set (basal polytomy + size ratio 30-70% + inter-lineage difference D=2 + MRCA placement) but conditioned I2 only on the squared polytomy sub-property (per O-22), so imposing the omitted size-ratio and sequence-difference conditions on I2 as well shrinks P(tau|I2) enough to reverse the Bayes factor.

Evaluation of the core inference. The symmetry principle is bedrock Bayesian practice — the likelihood must be evaluated on the same event for both hypotheses, and the burglary analogy (“blue Toyota” vs “drives a car”) is apt. Conditional on O-22 (that Pekar squared only the polytomy criterion), the fix is correct: under I2 the two independent introductions must jointly reproduce the observed size ratio and D=2, neither of which is automatic, so each contributes a multiplicative factor <1 that Pekar’s I2 likelihood omitted. The size-ratio factor is read off Pekar’s own Fig S22 via a growth-rate conversion (a ~4-day completion-time gap maps to |ln SR| ln(7/3) = 0.85 at an early-pandemic growth rate ~0.2/day — internally consistent), and the sequence-difference factor uses a Poisson D whose maximum P(D=2) = 2/e^2 = 0.271, marginalising to ~0.193. Both sub-steps are elementary and checkable; the multiplicative combination assumes rough independence of the polytomy, size-ratio and mutational factors, which is acceptable. Probed defeater: could Pekar’s asymmetry be justified because D=2 is set by reservoir diversity under I2 and thus “free”? No — even the most favourable tuning (E(D)=2) caps P(D=2) at 0.271, so I2 still pays a genuine penalty; the fix does not double-count, since both sides condition on D=2. The core inference (reversal to favour a single introduction) is valid.

Why corrected rather than approved. I traced the arithmetic and the reported magnitude does not hold as stated. The size-ratio fraction is given as “~0.162,” but the stated pair counts are 24,818 of 136,503, which equal 0.1818, not 0.162. With the internally consistent 0.1818, P(tau|I2) = 0.475^2 x 0.1818 x 0.193 = 0.0079 and the reversal factor is 0.031/0.0079 = 3.9, not 4.4; only the (unexplained, more favourable) 0.162 yields exactly 4.4. Since the writeup states its parameters are “chosen to favour I2” (i.e. conservative against the reversal), the defensible claim is a factor of order 4 (~3.9-4.4), not “at least 4.4.” Separately, the two presented routes are not reconciled: multiplying Pekar’s post-erratum BF 4.3 by 0.162 x 0.193 gives 0.134 (factor ~7.4), while the rebuild from polytomy^2 = 0.2256 gives factor 4.4 — the rebuild uses a higher I2 base (0.2256 vs Pekar’s implied 0.133), which is why it is the more conservative and the one reported. The direction of the reversal is robust across these choices; the precise “at least 4.4” is not, so statement is corrected to “a factor of order 4 (approximately 3.9-4.4).” Weissman’s restriction of the claim to Pekar’s specific A/B two-introduction result (not the broader origin question) is preserved.

Reasoning

A Bayes update multiplies the prior odds by P(data|I2)/P(data|I1), and the SAME observed properties of the data must be conditioned on for both hypotheses. P2022 instead conditioned I1 on a detailed property set (basal polytomy + size ratio 30-70% + sequence difference D=2 with no intermediates + MRCA placement) but conditioned I2 only on the polytomy sub-property, squared. Weissman’s analogy: updating burglary-suspect odds on ‘drives a blue Toyota’ for suspect 1 but only ‘drives a car’ for suspect 2 spuriously favours suspect 2, because the coarser property is more probable. The fix is to impose the same conditions on I2.

Size-ratio correction. Because I2’s two lineages are independent single introductions, the distribution of their size ratio can be read off P2022’s own I1 simulations (its Fig S22): simulations that take longer to reach a fixed case count are smaller at a fixed time, so times-to-35,000-cases convert to log-sizes. Pairing the 523 polytomy-meeting simulations, only pairs whose completion times differ by roughly 4 days fall within |ln SR|ln(7/3)=0.85. Of 523*261=136,503 non-self pairs, 24,818 differ by 4 days, giving a fraction ~0.162 that meets the size condition (a coarse read of Fig S22 alone already implies ~15-20%). Applying 0.162 to the 4.3 Bayes factor drops it to ~0.70.

Sequence-difference correction. I2 must also reproduce the observed inter-lineage difference D=2, and the proper Bayesian condition is P(D=2), not a one-sided extension. The inter-lineage difference is a sum of many independent low-probability mutations, so D is approximately Poisson; the maximum P(D=2|I2)=2/e^2=0.271 requires fine-tuning E(D)=2, and marginalising over a plausible prior on E(D) gives P(D=2|I2)=0.193 (including post-introduction pre-clade-root mutations changes this negligibly, to ~0.176).

Combination. Corrected P(tau|I2) 0.475^2 * 0.162 * 0.193 ~= 0.0071 (using parameters chosen to favour I2), versus P(tau|I1)=0.031, so the likelihood ratio is ~0.0071/0.031=0.23, i.e. at least 4.4 favouring a single introduction - a reversal of P2022’s conclusion. A piece of the size-ratio correction large enough to sink the two-introduction claim is visible in P2022’s own Figure S22 without accessing raw files. Weissman notes the 2024 Science erratum fixed only the three coding errors, leaving this deeper conditionalization error in the version of record, and that McCowan reached the same reversal by a different route (new balanced-framework simulations). Weissman restricts the claim to P2022’s specific A/B two-introduction result, not the broader origin question.

Link to original

Conclusion

H-16 - SARS-CoV-2 early lineages A and B are better explained by a single introduction than by two

Weissman stresses this bears only on P2022’s specific two-introduction claim and is not itself proof of a research-related origin.

Link to original