基于生化指标和机器学习建立疾病临床预测模型的研究进展

1邓梦洁,1,2,3 曾 俊,1郑力夫,1,2,3 王 璐,1,2,3 江 华

肿瘤代谢与营养电子杂志 ›› 2025, Vol. 12 ›› Issue (2) : 253-260.

PDF(980 KB)
PDF(980 KB)
肿瘤代谢与营养电子杂志 ›› 2025, Vol. 12 ›› Issue (2) : 253-260.
综述

基于生化指标和机器学习建立疾病临床预测模型的研究进展

  • 1邓梦洁,1,2,3 曾 俊,1郑力夫,1,2,3 王 璐,1,2,3 江 华
作者信息 +

Research progress in building clinical prediction models for diseases based on biochemical indicators and machine learning

  • 1Deng Mengjie, 1,2,3Zeng Jun, 1Zheng Lifu, 1,2,3Wang Lu, 1,2,3Jiang Hua
Author information +
文章历史 +

摘要

基于生化指标构建临床预测模型已成为现代医学研究的一个重要方向。 本文系统地梳理了近年来基于该模型的 研究进展,重点关注模型构建的关键环节,并指出此类模型目前存在的不足及未来的发展方向。 笔者通过检索 PubMed、Embase、Web of Science 和中国知网等数据库,对 28 项相关研究的建模目标、变量选择、建模方法及模型评价进行了总结和分析。 结果显示纳入的研究多为小样本、单中心设计,样本量中位数为 466 例(n = 54~ 58 616 例),研究对象以肿瘤患者为主,研究重 点集中于疾病预后的风险预测。 纳入研究的模型预测精度存在较大差异[曲线下面积(AUC)范围:0. 691~ 0. 992]。 进一步分 析显示,部分研究存在若干不足之处,包括纳入/ 排除标准不明确、数据预处理不规范、特征工程缺乏、未进行交叉验证和外部 验证等。 因此,尽管已有诸多研究尝试基于生化指标构建临床预测模型,但整体研究质量尚待提升。 未来研究应注重选择多 元自变量和先进机器学习算法进一步优化模型,并采用规范化的评估方法,确保模型的临床适用性,并重视时间序列数据的 应用,以构建高质量的预测模型,从而充分发挥生化指标的临床价值。

Abstract

The construction of clinical prediction models based on biochemical indicators has become an important direction in medical research. This paper systematically reviews recent advances in the development of clinical prediction models based on biochemical indicators and machine learning techniques with a focus on the key aspects of model construction highlighting the current limitations of such models and proposing future directions for improvement. We enrolled 28 studies that are retrieved from PubMed Embase Web of Science and China Knowledge Network CNKI databases summarizing and analyzing the modeling objectives variable selection modeling methods and model evaluations of these studies. We found that most of the included studies were small-sample single-center designs with a median sample size of 466 n = 54-58 616 . The study populations primarily consisted of cancer patients and the purpose of the studies was mainly prognosis/ risk prediction of disease. The model performance varied widely with AUC values ranging from 0. 691 to 0. 992 and the qualities of enrolled studies varied. Further analysis revealed several limitations including unclear inclusion / exclusion criteria lack of reliable preprocessing methods absence of feature engineering and insufficient cross-validation and external validation. Therefore although a few studies attempted to establish prediction models using biochemical indicators the overall quality of the research still needs improvement. Future research should focus on optimizing models using multivariate variable selection and advanced machine learning / deep learning algorithms adopting standardized evaluation methods for model validation to ensure the clinical applicability of the models and incorporating time-series data to enhance model quality and fully realize the clinical value of biochemical indicators.

关键词

生化指标 / 代谢指标 / 预测模型 / Logistic 回归 / 机器学习 / 人工智能 / 联合建模 / 数字孪生

Key words

Biochemical indicators / Metabolic indicators / Predictive modeling / Logistic regression / Machine learning / Artificial intelligence / Combined modeling / Digital twin

引用本文

导出引用
1邓梦洁,1,2,3 曾 俊,1郑力夫,1,2,3 王 璐,1,2,3 江 华. 基于生化指标和机器学习建立疾病临床预测模型的研究进展[J]. 肿瘤代谢与营养电子杂志. 2025, 12(2): 253-260
1Deng Mengjie, 1,2,3Zeng Jun, 1Zheng Lifu, 1,2,3Wang Lu, 1,2,3Jiang Hua. Research progress in building clinical prediction models for diseases based on biochemical indicators and machine learning[J]. Electronic Journal of Metabolism and Nutrition of Cancer. 2025, 12(2): 253-260

基金

四川省科学技术厅重点项目资助(2024YFFK0075)

PDF(980 KB)

Accesses

Citation

Detail

段落导航
相关文章

/