What the analysis says
The cluster asks what causal role the Huanan market played: the primary wildlife-to-human spillover origin (H-14), an externally-seeded amplifier that took in a virus introduced from outside by infected people or cold-chain goods (H-15), a detection/ascertainment artifact that was not causally special (H-19), or something unlisted (H-45, residual). The prior was near-balanced across the three named roles — model prior [0.3826, 0.3532, 0.2075, 0.0566] — an outside view on venue-cluster base rates (SARS-CoV-1 markets pull toward origin, Ebola/MERS/Nipah venues toward amplifier) that deliberately left the origin-vs-amplifier split a near-coin-flip (share_origin = 0.52) for the Wuhan evidence to resolve.
That evidence updated hard toward the market as spillover origin: posterior [0.7401, 0.1842, 0.0483, 0.0275]. Two items did most of the work against the artifact reading. The geospatial pattern (CG-7, on dataset D-2) — cases clustering on the market against a population-density null, both lineages centering there, market-unlinked cases living even closer than linked ones (4.00 vs 5.74 km) — is ~2.5× better explained by a causally-central market than by a pure artifact (lik 1.0 vs 0.4), and the market-root tMRCA with both basal lineages present (E-19) adds a signature a downstream/incidental venue does not predict (H-14 vs H-19 lik 1.0 vs 0.3). Together they pushed H-19 down to ~5%. The environmental-metagenomic pattern (CG-8, on China-CDC dataset D-1) — viral positives concentrated ~87.5% in the wildlife wing and at the one live-mammal stall, co-located with susceptible-mammal genetic material — is what gives H-14 its edge over H-15, but only a ~1.4× one (lik 1.0 vs 0.7). The pre-pandemic wildlife-trade inventory (CG-10) establishes the animal-side precondition, not the event, so it barely moved the vector.
What the model may not capture
The entire object-level zoonosis signal rests on two shared, single-witness datasets: D-2 under all of CG-7, and D-1 under all twelve observations in CG-8. Both were collected by parties with agenda concerns — D-2 by field teams who actively sought market links and concentrated on central-Wuhan hospitals near the market, D-1 by China CDC as sole physical actor, released late and only partially. Both have been re-analysed to opposite conclusions on the same bytes. This is a motivated-source, single-witness structure: if either collection is compromised, every observation on it moves together. The joint-likelihood treatment (correctly) prices each dataset as one witness, not six or twelve independent draws, but cannot rule out that the witness is wrong.
The residential-proximity ascertainment bias is a live escape H-19 could still exploit. A-44 rebuts market-link ascertainment (unlinked cases cluster too) but concedes a distinct residual bias — case-finding concentrated at central hospitals near the market — that could inflate even the unlinked-cases-closer datum the edge rests on. CG-7 docks trust for exactly this (t=0.50), but not fully; if it dominates, H-19 is undervalued.
Is the answer on the list? The residual holds only ~3%, but the trichotomy assumes a clean single-type role. It cannot express “partly both” — a genuine animal spillover and substantial human amplification at the same venue — which much of the CG-8 pattern fits and the mutually-exclusive members forbid. An unlisted variant like cold-chain-as-true-origin (folded into H-45) would still make the market causally special, so it is about as consequential here as the two live members; the shared framing that the role is one clean type is the standing doubt, and 3% likely under-weights it.
What would help
- Independent sampling of the market’s live-wildlife stalls from before closure and animal removal — does not exist. D-1’s sampling began 1 Jan 2020 after the animals were cleared, so there is no direct test of whether a market animal was infected; this is the single missing datum the whole H-14-vs-H-15 crux turns on.
- The un-released China-CDC collection metadata behind D-1 (full sampling frame, per-stall qPCR status, the withheld raw material) — exists, inaccessible.
- An early case line-list free of market-seeking ascertainment — does not exist; the bias is baked into D-2 at collection, so no reanalysis can fully remove it.
Confusions and contradictions
The H-14-vs-H-15 edge — the fragile part of this cluster — rests on the single most contestable number in the analysis (CG-8’s ~1.4×), which sits on an irreducible conflict between two competent re-analyses of the same dataset D-1. Crits-Christoph (A-48, A-49) read the stall-A co-location of viral RNA with susceptible-mammal mtDNA as a wildlife-source signal and the whole-dataset negative correlation as a sampling-design + timing artifact (their balanced n=70 set: human mtDNA uncorrelated, porcupine/marmot positive). Bloom reads the same data the opposite way (A-54): viral abundance tracks fish/livestock/human material and is negatively associated with raccoon dog (13 of 14 raccoon-dog-dominated samples had zero reads), the pattern of human deposition — and, cutting both ways (A-55), samples taken ~1 month into widespread human transmission make co-location a weak, non-decisive instrument in either direction. The analysis lands H-14 above H-15 but flags this as contested, not resolved; on current data it is not adjudicable. Shipped unfixed, correctly.
External consensus
The mainstream epidemiological/virological position (Worobey, Pekar, Crits-Christoph, the WHO SAGO) treats the market as the most likely site of emergence and the early-case geography as real signal, while a vocal minority (Bloom’s metagenomic reanalysis, Stoyan-Chiu’s spatial critique, Weissman) holds it non-dispositive. The posterior’s ~74% for origin sits within that zoonosis-leaning majority but is more resolved on origin-vs-amplifier specifically than the D-1 dispute above arguably licenses — the divergence is at the H-14/H-15 boundary, not on rejecting the artifact reading, where the analysis and the field agree.