Annotating the microbial dark matter with HiFi-NN

利用HiFi-NN对微生物暗物质进行注释

阅读:1

Abstract

The accurate computational annotation of protein sequences with enzymatic function remains a fundamental challenge in bioinformatics. Here, we present HiFi-NN (Hierarchically-Finetuned Nearest Neighbor search) which annotates protein sequences to the 4th level of Enzyme Commission (EC) number with greater precision and recall than state-of-the-art deep learning methods. Furthermore, we show that this method can correctly identify the EC number of a given sequence to lower identities than BLASTp. We show that performance can be improved by increasing the diversity of the lookup set in both sequence space and the environment the sequence has been sampled from. We proceed to show that we can correct specific mis-annotations in the BRENDA enzymes database reproducing results found by others. Finally, we use HiFi-NN to annotate functional dark-matter protein sequences from NMPFamDB. Our findings pave the way for more accurate functional annotation in silico, especially for proteins from distant sequence space.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。