Early Detection of Novel SARS-CoV-2 Variants from Urban and Rural Wastewater through Genome Sequencing and Machine Learning

利用基因组测序和机器学习技术早期检测城乡污水中的新型SARS-CoV-2变异株

阅读:1

Abstract

Genome sequencing from wastewater has emerged as an accurate and cost-effective tool for identifying SARS-CoV-2 variants. However, existing methods for analyzing wastewater sequencing data are not designed to detect novel variants that have not been characterized in humans. Here, we present an unsupervised learning approach that clusters co-varying and time-evolving mutation patterns leading to the identification of SARS-CoV-2 variants. To build our model, we sequenced 3,659 wastewater samples collected over a span of more than two years from urban and rural locations in Southern Nevada. We then developed a multivariate independent component analysis (ICA)-based pipeline to transform mutation frequencies into independent sources with co-varying and time-evolving patterns and compared variant predictions to >5,000 SARS-CoV-2 clinical genomes isolated from Nevadans. Using the source patterns as data-driven reference "barcodes", we demonstrated the model's accuracy by successfully detecting the Delta variant in late 2021, Omicron variants in 2022, and emerging recombinant XBB variants in 2023. Our approach revealed the spatial and temporal dynamics of variants in both urban and rural regions; achieved earlier detection of most variants compared to other computational tools; and uncovered unique co-varying mutation patterns not associated with any known variant. The multivariate nature of our pipeline boosts statistical power and can support accurate and early detection of SARS-CoV-2 variants. This feature offers a unique opportunity for novel variant and pathogen detection, even in the absence of clinical testing.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。