Development, validation and recalibration of a prediction model for prediabetes: an EHR and NHANES-based study

基于电子健康记录和美国国家健康与营养调查(NHANES)的糖尿病前期预测模型的开发、验证和重新校准研究

阅读:1

Abstract

BACKGROUND: A prediction model that estimates the risk of elevated glycated hemoglobin (HbA1c) was developed from electronic health record (EHR) data to identify adult patients at risk for prediabetes who may otherwise go undetected. We aimed to assess the internal performance of a new penalized regression model using the same EHR data and compare it to the previously developed stepdown approximation for predicting HbA1c ≥ 5.7%, the cut-off for prediabetes. Additionally, we sought to externally validate and recalibrate the approximation model using 2017-2020 pre-pandemic National Health and Nutrition Examination Survey (NHANES) data. METHODS: We developed logistic regression models using EHR data through two approaches: the Least Absolute Shrinkage and Selection Operator (LASSO) and stepdown approximation. Internal validation was performed using the bootstrap method, with internal performance evaluated by the Brier score, C-statistic, calibration intercept and slope, and the integrated calibration index. We externally validated the approximation model by applying original model coefficients to NHANES, and we examined the approximation model's performance after recalibration in NHANES. RESULTS: The EHR cohort included 22,635 patients, with 26% identified as having prediabetes. Both the LASSO and approximation models demonstrated similar discrimination in the EHR cohort, with optimism-corrected C-statistics of 0.760 and 0.763, respectively. The LASSO model included 23 predictor variables, while the approximation model contained 8. Among the 2,348 NHANES participants who met the inclusion criteria, 30.1% had prediabetes. External validation of the LASSO model was not possible due to the unavailability of some predictor variables. The approximation model discriminated well in the NHANES dataset, achieving a C-statistic of 0.787. CONCLUSION: The approximation method demonstrated comparable performance to LASSO in the EHR development cohort, making it a viable option for healthcare organizations with limited resources to collect a comprehensive set of candidate predictor variables. NHANES data may be suitable for externally validating a clinical prediction model developed with EHR data to assess generalizability to a nationally representative sample, depending on the model's intended use and the alignment of predictor variable definitions with those used in the model's original development.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。