A systematic assessment of machine learning for structural variant filtering

对用于结构变异过滤的机器学习进行系统评估

阅读：2

作者：Kalra,Archit,Paulin,Luis F,Sedlazeck,Fritz J

期刊：		影响因子：
时间：	2026	起止号：	2026 Jan 30
doi：	10.64898/2026.01.27.702059

Abstract

BACKGROUND: Accurate discrimination of true structural variants (SVs) from artifacts in long-read sequencing data remains a critical bottleneck. Numerous machine learning solutions have been proposed, ranging from classical models using engineered features to advanced deep learning and foundation model interpretability methods. However, a systematic comparison of their performance, efficiency, and practical utility is lacking. RESULTS: We conducted a comprehensive benchmark of five machine learning paradigms for SV filtering using standardized Genome in a Bottle (GIAB) data for samples HG002 and HG005. We evaluated classical Random Forest classifiers on 15 genomic features, computer vision models (ResNet/VICReg), diffusion-based anomaly detection, sparse autoencoders (SAEs) on the Evo2-7B foundation model, and multimodal ensembles. A simple Random Forest on interpretable features achieved a peak F1-score of 95.7%, effectively matching all more complex models (ResNet50: 95.9%, Diffusion: 95.8%). This study represents the first application of diffusion-based anomaly detection and sparse autoencoders to structural variant analysis; while diffusion models learned highly discriminative, disentangled representations and SAEs uncovered biologically interpretable features (including atoms that were specific for ALU deletions, chromosome X variants and insertion events), they did not significantly surpass this classification ceiling. Ensemble methods offered no performance benefit but may have future potential given the orthogonality of vision-based and linear features. CONCLUSIONS: Our findings demonstrate that for the established task of germline SV filtering, simpler, interpretable models provide an optimal balance of accuracy, speed, and transparency. This benchmark establishes a pragmatic framework for method selection and argues that increased model complexity must be justified by clear, unmet biological needs rather than marginal predictive gains.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用；引用内容仅为补充信息，不代表本站立场。

2、若认为本页面引用内容涉及侵权，请及时与本站联系，我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容，需注明“来源：[生知库]”并获得授权；使用引用内容的，需自行联系原作者获得许可。

4、投稿及合作请联系：info@biocloudy.com。

肿瘤免疫

炎症

T细胞

凋亡

线粒体

转录调控

巨噬细胞

自噬

传染病

氧化应激

血管生成

磷酸化

肠道菌群

囊泡

3D/类器官

中性粒细胞

单细胞

外泌体

药物研究

DNA甲基化

细胞衰老

miRNA

铁死亡

缺氧低氧

乙酰化

泛素化

组蛋白修饰

炎性小体

代谢重编程

树突状细胞

焦亡

肿瘤微环境

lncRNA

m6A/m5C/m7G

空间多组学

细胞基因治疗

内质网应激

相分离

治疗耐药

Treg

免疫代谢

上皮间质转化

染色质重塑

脂质过氧化

蛋白质稳态

铁代谢

脂代谢

cGAS-STING

肠脑轴

细胞极性

碱基编辑

氨基酸代谢

乳酸化

蛋白降解

翻译调控

circRNA

低氧缺氧

piRNA

肿瘤异质性

NK 细胞

氧化脂质

MDSC

溶酶体功能

NETosis

RNA 编辑

细胞干性

琥珀酰化

CAR-NK

冷应激

器官芯片

Tfh

巴豆酰化

表观遗传记忆

空间代谢组

器官纤维化

线粒体未折叠蛋白反应

铜死亡

自噬流

程序性坏死

肠肝轴

MAIT 细胞

丙酰化