summary - A Wuhan University team had deposited, then in June 2020 requested deletion of, 241 sequencing runs (project PRJNA612766) from NCBI’s Sequence Read Archive; NCBI deleted rather than merely hid the files. Bloom recovered the deleted files from Google Cloud storage (by reconstructing archived download URLs) and reconstructed partial genomes for 45 early-epidemic samples collected in Wuhan in January 2020. These recovered sequences carried the C29095T mutation and were less likely to carry the T8782C/C28144T mutation pair (which defines the market-linked “lineage B”) than the sequences reported in the WHO-China joint report — i.e., they were genetically closer to bat-coronavirus relatives than the heavily-studied market-associated samples. Bloom argues this shows the published/market-associated sequence set does not capture the full viral diversity that was already circulating in Wuhan in January 2020, i.e. an ascertainment bias toward market-associated viruses in what got sequenced, sampled, and reported.
relevance_note - The core ascertainment-bias critique of the market-centric case/sequence timeline that both Worobey 2022 and the WHO report rely on: argues the “earliest known cases/sequences cluster at the market” pattern may partly reflect what got reported and retained, not the true index-case pattern. A 2025 bioRxiv/MBE reexamination (not separately noded, since it reanalyzes these same recovered sequences) pushes back, arguing the recovered sequences’ collection date (30 January 2020) was actually later than hundreds of other published sequences and that Wuhan-exposure history was common across early samples generally, so the single family cluster Bloom highlights doesn’t establish broader undersampling.