PRIMED: predicting DNA binding residues by leveraging pre-trained protein language models

PRIMED：利用预训练的蛋白质语言模型预测DNA结合残基

阅读：1

作者：Zhang,Luoshu,Li,Xin,Song,Ruocen,Song,Qianqian,Fan,Xiao

期刊：	Frontiers in Artificial Intelligence	影响因子：	4.700
时间：	2026	起止号：	2026;9:1763313
doi：	10.3389/frai.2026.1763313

Abstract

INTRODUCTION: Protein-DNA interactions are central to gene regulation, genome stability, and disease mechanisms. Identifying DNA-binding residues (DBRs) is critical for structural modeling, protein engineering, and therapeutic design. Although experimental approaches provide valuable insights, they remain low-throughput and resource-intensive. Computational methods offer scalable alternatives by leveraging protein sequential and structural information to predict DBRs. METHODS: We present PRIMED (Protein Residue Inference using Multilayer perceptron for Enhanced DNA-binding predictions), a machine learning framework that integrates protein representations of distinct biochemical and structural properties from three protein language models: ESM-2, ESM-3, and ESM-C. These representations are concatenated and processed by a multilayer perceptron to perform DBR predictions. RESULTS: PRIMED demonstrated strong performance across three benchmark datasets: Test-46 and Test-129 from a previous study, CLAPE-DB, and Test-10 K, which we curated from UniProtKB/Swiss-Prot. The model achieves an area under the Receiver Operating Characteristic curve (AUC) of 0.92 and a Matthews Correlation Coefficient (MCC) of 0.64 on Test-46, as well as an AUC of 0.93 and MCC of 0.45 on Test-129. On Test-10 K, PRIMED demonstrates generalizability across proteins with varying DBR percentages, maintaining competitive performance relative to the runner-up method, CLAPE-DB. DISCUSSION: These results highlight the effectiveness of integrating diverse protein language model representations for accurate, transferable DBR predictions.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用；引用内容仅为补充信息，不代表本站立场。

2、若认为本页面引用内容涉及侵权，请及时与本站联系，我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容，需注明“来源：[生知库]”并获得授权；使用引用内容的，需自行联系原作者获得许可。

4、投稿及合作请联系：info@biocloudy.com。

肿瘤免疫

炎症

T细胞

线粒体

凋亡

转录调控

巨噬细胞

自噬

传染病

氧化应激

肠道菌群

磷酸化

血管生成

囊泡

3D/类器官

单细胞

中性粒细胞

外泌体

DNA甲基化

miRNA

药物研究

铁死亡

细胞衰老

乙酰化

缺氧低氧

泛素化

树突状细胞

组蛋白修饰

炎性小体

肿瘤微环境

lncRNA

代谢重编程

焦亡

m6A/m5C/m7G

内质网应激

空间多组学

细胞基因治疗

治疗耐药

相分离

Treg

上皮间质转化

免疫代谢

染色质重塑

脂质过氧化

脂代谢

蛋白质稳态

铁代谢

细胞极性

氨基酸代谢

碱基编辑

cGAS-STING

肠脑轴

蛋白降解

乳酸化

翻译调控

circRNA

piRNA

肿瘤异质性

NK 细胞

氧化脂质

MDSC

NETosis

低氧缺氧

溶酶体功能

细胞干性

琥珀酰化

CAR-NK

冷应激

RNA 编辑

Tfh

巴豆酰化

器官芯片

表观遗传记忆

铜死亡

器官纤维化

线粒体未折叠蛋白反应

空间代谢组

程序性坏死

自噬流

肠肝轴

丙酰化

MAIT 细胞