Effects of the training dataset characteristics on the performance of nine species distribution models: application to Diabrotica virgifera virgifera

训练数据集特征对九种物种分布模型性能的影响：以Diabrotica virgifera virgifera为例

阅读：1

作者：Dupin,Maxime,Reynaud,Philippe,Jarošík,Vojtěch,Baker,Richard,Brunel,Sarah,Eyre,Dominic,Pergl,Jan,Makowski,David

期刊：	PLoS One	影响因子：	2.600
时间：	2011	起止号：	2011;6(6):e20957
doi：	10.1371/journal.pone.0020957

Abstract

Many distribution models developed to predict the presence/absence of invasive alien species need to be fitted to a training dataset before practical use. The training dataset is characterized by the number of recorded presences/absences and by their geographical locations. The aim of this paper is to study the effect of the training dataset characteristics on model performance and to compare the relative importance of three factors influencing model predictive capability; size of training dataset, stage of the biological invasion, and choice of input variables. Nine models were assessed for their ability to predict the distribution of the western corn rootworm, Diabrotica virgifera virgifera, a major pest of corn in North America that has recently invaded Europe. Twenty-six training datasets of various sizes (from 10 to 428 presence records) corresponding to two different stages of invasion (1955 and 1980) and three sets of input bioclimatic variables (19 variables, six variables selected using information on insect biology, and three linear combinations of 19 variables derived from Principal Component Analysis) were considered. The models were fitted to each training dataset in turn and their performance was assessed using independent data from North America and Europe. The models were ranked according to the area under the Receiver Operating Characteristic curve and the likelihood ratio. Model performance was highly sensitive to the geographical area used for calibration; most of the models performed poorly when fitted to a restricted area corresponding to an early stage of the invasion. Our results also showed that Principal Component Analysis was useful in reducing the number of model input variables for the models that performed poorly with 19 input variables. DOMAIN, Environmental Distance, MAXENT, and Envelope Score were the most accurate models but all the models tested in this study led to a substantial rate of mis-classification.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用；引用内容仅为补充信息，不代表本站立场。

2、若认为本页面引用内容涉及侵权，请及时与本站联系，我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容，需注明“来源：[生知库]”并获得授权；使用引用内容的，需自行联系原作者获得许可。

4、投稿及合作请联系：info@biocloudy.com。

肿瘤免疫

炎症

T细胞

线粒体

凋亡

转录调控

巨噬细胞

自噬

传染病

氧化应激

肠道菌群

磷酸化

囊泡

血管生成

3D/类器官

单细胞

中性粒细胞

外泌体

DNA甲基化

miRNA

药物研究

铁死亡

细胞衰老

乙酰化

缺氧低氧

泛素化

树突状细胞

炎性小体

肿瘤微环境

组蛋白修饰

lncRNA

代谢重编程

焦亡

m6A/m5C/m7G

内质网应激

空间多组学

细胞基因治疗

相分离

治疗耐药

Treg

上皮间质转化

免疫代谢

染色质重塑

脂质过氧化

蛋白质稳态

脂代谢

铁代谢

细胞极性

氨基酸代谢

cGAS-STING

碱基编辑

蛋白降解

肠脑轴

翻译调控

乳酸化

circRNA

piRNA

肿瘤异质性

NK 细胞

氧化脂质

MDSC

NETosis

溶酶体功能

低氧缺氧

琥珀酰化

细胞干性

CAR-NK

冷应激

RNA 编辑

Tfh

巴豆酰化

器官芯片

器官纤维化

表观遗传记忆

铜死亡

线粒体未折叠蛋白反应

空间代谢组

程序性坏死

自噬流

丙酰化

MAIT 细胞

肠肝轴