Comprehensive duck DNA fingerprinting based on machine learning for breed identification

基于机器学习的鸭子DNA指纹图谱技术在品种鉴定中的应用

阅读:2

Abstract

Duck is one of the most widely distributed waterfowl in the world, with more than 6 billion of them farmed annually in the world, and has great economic and ecological value. Amidst mounting global prioritization of duck genetic resource exploration and prevalent inter-varietal hybridization events, the traditional breed identification methods are difficult to address actual requirements, restricting the utilization, development and protection of duck germplasm resources. This study aims to develop an accurate, efficient, and scalable duck DNA fingerprinting system based on genomic technologies and machine learning methods to address the urgent need for breed identification tools in high-quality agricultural production and ecological protection. Our study aims to construct a global duck DNA fingerprint map based on genomic data and machine learning algorithm, develop an accurate, efficient and scalable duck DNA fingerprinting identification tool, and solve the urgent need for breed identification tools for high-quality agricultural production and ecological protection. In this study, we obtained the whole genome resequencing data of 196 duck individuals from 16 breeds and constructed a high-density duck population variation dataset containing 2,360,039 SNPs. Four characteristic molecular marker selection methods (Delta, Average Euclidean Distance (AED), Polymorphism Information Content (PIC), and Fixation Index (F(ST))) and four machine learning classification algorithms (Random Forest (RF), Support Vector Machine (SVM), Linear Discriminant Analysis (LDA), and Naive Bayes (NB)) were tested. The results showed that AED indicator had the best performance in selecting SNP markers in ducks, and the classification accuracy was the highest (98.38 %) when 2000 SNP sites were selected. SVM algorithm showed the best classification performance in ducks, with the classification accuracy of 98.71 % and the running time was within 70 seconds. We constructed the duck DNA fingerprinting maps of 16 breeds based on the AED indicator and SVM algorithm, each containing 200 SNP markers. We have also developed a user-friendly and efficient duck DNA fingerprinting identification tool that could achieve identification of large-scale genetic resources, and also collect new duck genetic resources and use them for breed identification. Our results provide advanced method and utility tool support for identifying and utilizing world-wide duck germplasm resources and a reference for the development of DNA fingerprinting maps for other major agricultural animals.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。