Logistic回归法和机器学习算法构建子痫前期预测模型的比较
Comparison of logistic regression and machine learning algorithm in establishment of pre-eclampsia prediction model
摘要目的:利用医院电子病历系统首页信息和临床检验数据通过logistic回归(logistic regression,LR)法和机器学习算法构建子痫前期(preeclampsia,PE)预测模型,同时比较机器学习算法和LR构建模型的预测性能。方法:基于2012年1月1日至2019年12月31日在广州医科大学附属第三医院就诊孕产妇的围产期数据和柔济妊娠检验数据库信息,根据临床诊疗指南和相关文献报道关联整合后选取数据量较为完整的孕24~28周共2 736例孕妇的28项临床相关指标作为PE预测模型构建数据。将其中PE患者作为PE组( n=245),其余非PE患者中采用欠采样法选择255例为对照组。使用随机森林算法(random forest,RF)及极端梯度上升算法(eXtreme Gradient Boosting,XGB)和LR模型分别构建PE疾病预测模型。模型构建完成后,在2019年6月至2022年12月开展的PE前瞻性队列研究获得数据中(PE组38例,对照组80例),进行PE预测准确性的外部验证。采用准确度、灵敏度、特异度、受试者工作特征曲线下面积比较不同模型的预测效能。 结果:3种预测模型构建时纳入的指标提示,尿酸、肌酐、年龄、孕早期体重指数、尿素、甘油三酯、红细胞计数、嗜酸性粒细胞计数、总胆固醇、中性粒细胞计数、尿蛋白、丙氨酸氨基转移酶以及尿潜血是影响PE预测模型的指标。RF、XGB和LR模型在训练集和测试集中的受试者工作特征曲线下面积分别为0.851(95% CI:0.730~0.891)、0.955(95% CI:0.865~0.987)、0.884(95% CI:0.767~0.923)和0.845(95% CI:0.723~0.868)、0.907(95% CI:0.791~0.919)、0.851(95% CI:0.755~0.893)。在测试集中,RF、XGB和LR模型的准确度、灵敏度与特异度分别为0.803、0.607、0.958,0.864、0.790、0.927和0.832、0.661、0.971。在外部验证集中RF、XGB和LR预测模型的准确度分别为0.822、0.814和0.763;灵敏度分别为0.737、0.789和0.605;特异度分别为0.863、0.825和0.838,其中XGB模型的约登指数最高,为0.614。 结论:相对于传统的建模方法,利用机器学习算法可以在真实临床检测数据中建立更加有效的PE预测模型。
更多相关知识
abstractsObjective:To construct preeclampsia (PE) prediction models using information from the hospital electronic medical information and clinical laboratory data through logistic regression (LR) and machine learning algorithms, and to compare their predictive performance.Methods:The study was conducted based on the information from Rouji Pregnancy Test Database and the perinatal data of women who visited the Third Affiliated Hospital of Guangzhou Medical University from January 1, 2012, to December 31, 2019. Drawing upon clinical treatment guidelines and related literature, 28 clinical indicators from 2 736 pregnant women at 24 to 28 weeks of gestation were selected after a thorough integration and used for the construction of the PE prediction model dataset. Patients diagnosed with PE comprised the PE group ( n=245), while another 255 cases from the rest who did not have PE were selected, with undersampling method, as the control group. The Random Forest algorithm (RF), eXtreme Gradient Boosting (XGB) algorithm, and LR model were each employed to develop predictive models for PE. Following the construction of the models, external validation of PE prediction accuracy was carried out using data acquired from an independent prospective cohort study on PE that was conducted from June 2019 to December 2022, in which 38 PE cases and 80 controls were chosen. The performance of predictive models were evaluated using metrics such as accuracy, sensitivity, specificity, and the area under the curve (AUC) of receiver operating characteristic. Results:Indicators included in the construction of the three predictive models suggested that uric acid, creatinine, maternal age, early pregnancy body mass index, urea, triglycerides, red blood cell count, eosinophil count, total cholesterol, neutrophil count, urine protein, alanine aminotransferase, and urine occult blood were influential in PE prediction models. The AUCs for RF, XGB, and LR models in the training and test sets were 0.851 (95% CI:0.730-0.891), 0.955 (95% CI:0.865-0.987), 0.884 (95% CI:0.767-0.923) vs. 0.845 (95% CI:0.723-0.868), 0.907 (95% CI:0.791-0.919), 0.851 (95% CI:0.755-0.893), respectively. In the test set, the accuracy, sensitivity, and specificity for RF, XGB, and LR models were 0.803, 0.607, 0.958, 0.864, 0.790, 0.927, and 0.832, 0.661, 0.971, respectively. In the external validation of the RF, XGB and LR predictive models, the accuracy were 0.822, 0.814, and 0.763; the sensitivity were 0.737, 0.789, and 0.605, and the specificity were 0.863, 0.825, and 0.838, respectively. Among them, XGB model showed the highest Youden's index (0.614). Conclusion:Compared to traditional methods of model construction, machine learning algorithms can establish more effective PE prediction models using real clinical data.
More相关知识
- 浏览56
- 被引4
- 下载1

相似文献
- 中文期刊
- 外文期刊
- 学位论文
- 会议论文


换一批



