Reports that a set of ~13 early-outbreak SARS-CoV-2 sequencing files were deleted from the NIH Sequence Read Archive in June 2020 at the depositing Chinese team’s request, recovers them from Google Cloud project caches, and partially reconstructs the underlying sequences. Argues the recovered sequences carry mutations placing them closer to the bat-coronavirus outgroup (RaTG13) than the market-associated genomes that dominate the public early record, and that this is consistent with (and improves the case for) an earlier, more bat-CoV-like progenitor genotype circulating in Wuhan before the market cluster — reinforcing Kumar et al.’s (S-3) progenitor argument from an independent data-recovery angle. Became a focal point of the lab-leak-side “the early record was scrubbed” narrative regardless of Bloom’s own more measured within-paper interpretation.

relevance_note: the primary data-recovery paper behind “the deleted-sequences” argument; whether its dating/representativeness claims hold up is now directly contested by S-11.