ML-ExonCNV: a robust XGBoost multi-expert ensemble framework for rare exon CNV detection in whole-exome sequencing data

ML-ExonCNV:一种用于全外显子测序数据中罕见外显子拷贝数变异检测的稳健的 XGBoost 多专家集成框架

阅读:1

Abstract

Copy number variants (CNVs) have been shown to play a significant role in the pathogenesis of various human diseases. Although several tools have been developed for detecting CNVs based on whole-exome sequencing (WES) data, their performance remains suboptimal for small exon-level CNVs (exCNVs). This is primarily due to multiple technical variabilities, including probe capture efficiency, mappability, exon size, batch effects, experimental background noise bias, and control sample selection, all of which can lead to false negatives and false positives in exCNV detection. To address these challenges, we developed ML-ExonCNV, which innovatively integrates the XGBoost machine learning model with a multi-expert ensemble approach. The model was trained using 14 features derived from 22 364 real-world, quantitative polymerase chain reaction-validated rare exCNVs. Evaluation on a test set of 492 real WES and the NA12878 gold-standard dataset demonstrated that ML-ExonCNV outperformed widely used tools such as GATK-gCNV, ExomeDepth, and CNVkit. Notably, ML-ExonCNV can detect large segmental CNVs, mosaic CNVs, and breakpoint CNVs on exon region. Furthermore, our analysis revealed recurrent exCNV-associated genes and their phenotypic correlations. Neurodevelopmental and musculoskeletal abnormalities were identified as the most frequently associated phenotypes with high-recurrence exCNVs.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。