Welcome to visit Zhongnan Medical Journal Press Series journal website!

PROBAST+AI: an interpretation of the tool for assessing the quality, risk of bias and applicability of prediction models based on traditional or artificial intelligence methods

Published on Aug. 04, 2026Total Views: 56 timesTotal Downloads: 12 timesDownloadMobile

Author: ZOU Wenting 1, 2 ZHENG Qicheng 1, 2 HUANG Qiao 1 WANG Yongbo 1 REN Xiangying 1 XIAO Shifeng 1, 2 JIN Yinghui 1 YAN Siyu 1

Affiliation: 1.Center for Evidence-Based and Translational Medicine, Zhongnan Hospital of Wuhan University, Wuhan 430071, China 2.The Second Clinical Medical College, Wuhan University, Wuhan 430071, China

Keywords: PROBAST+AI Prediction model Systematic review Risk of bias Artificial intelligence Machine learning

DOI: 10.12173/j.issn.1004-5511.202601172

Reference: Citation:Zou WT, Zheng QC, Huang Q, et al. PROBAST+AI: an interpretation of the tool for assessing the quality, risk of bias and applicability of prediction models based on traditional or artificial intelligence methods[J].Yixue Xinzhi Zazhi, 2026, 36(7): 721-733. DOI: 10.12173/j.issn.1004-5511.202601172.[Article in Chinese]

  • Abstract
  • Full-text
  • References
Abstract

In recent years, artificial intelligence (AI) or machine learning (ML) methods have been increasingly used in the development and evaluation of clinical prediction models. Their algorithms differ from traditional regression modeling methods, resulting in significant limitations of the existing Prediction Model Risk of Bias Assessment Tool (PROBAST) for their evaluation. To address these limitations, the tool is updated to PROBAST+AI in 2025. PROBAST+AI extends the original framework to specifically address methodological challenges unique to developing or evaluating clinical prediction models based on AI/ML algorithms. PROBAST+AI assesses the quality of model development and the risk of bias in model evaluation across 4 domains: participants and data sources, predictors, outcome, and analysis, encompassing 16 and 18 signaling questions, respectively. Furthermore, the tool evaluates the applicability of the model across 3 domains: participants and data sources, predictors, and outcome. This article aims to compare the changes between the original and updated versions of the PROBAST tool, interpret the key content and items of PROBAST+AI, and apply the updated tool to evaluate an example clinical prediction model publication, to help domestic systematic review authors, clinicians, and policymakers critically appraise studies that develop or evaluate prediction models based on traditional or AI/ML methods.

Full-text
Please download the PDF version to read the full text: download
References

1. DamenJAAG, HooftL, SchuitE, et al. Prediction models for cardiovascular disease risk in the general population: systematic review[J]. BMJ, 2016, 353: i2416. doi:10.1136/bmj.i2416

2. BellouV, BelbasisL, KonstantinidisAK, et al. Prognostic models for outcome prediction in patients with chronic obstructive pulmonary disease: systematic review and critical appraisal[J]. BMJ, 2019, 367: l5358. doi:10.1136/bmj.l5358

3. de MunterL, PolinderS, LansinkKW, et al. Mortality prediction models in the general trauma population: a systematic review[J]. Injury, 2017, 48(2): 221-229. doi:10.1016/j.injury.2016.12.009

4. WynantsL, Van CalsterB, CollinsGS, et al. Prediction models for diagnosis and prognosis of COVID-19: systematic review and critical appraisal[J]. BMJ, 2020, 369: m1328.

5. Seyyed-KalantariL, ZhangH, McDermottMBA, et al. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations[J]. Nat Med, 2021, 27(12): 2176-2182. doi:10.1038/s41591-021-01595-0

6. YangY, ZhangH, GichoyaJW, et al. The limits of fair medical imaging AI in real-world generalization[J]. Nat Med, 2024, 30(10): 2838-2848. doi:10.1038/s41591-024-03113-4

7. WolffRF, MoonsKGM, RileyRD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies[J]. Ann Intern Med, 2019, 170(1): 51-58. doi:10.7326/m18-1376

8. SchinkelM, ParanjapeK, Nannan PandayRS, et al. Clinical applications of artificial intelligence in sepsis: a narrative review[J]. Comput Biol Med, 2019, 115: 103488. doi:10.1016/j.compbiomed.2019.103488

9. HurkmansC, BibaultJE, ClementelE, et al. Assessment of bias in scoring of AI-based radiotherapy segmentation and planning studies using modified TRIPOD and PROBAST guidelines as an example[J]. Radiother Oncol, 2024, 194: 110196. doi:10.1016/j.radonc.2024.110196

10. CaiY, CaiYQ, TangLY, et al. Artificial intelligence in the risk prediction models of cardiovascular disease and development of an independent validation screening tool: a systematic review[J]. BMC Med, 2024, 22(1): 56. doi:10.1186/s12916-024-03273-7

11. KaulT, DamenJA, WynantsL, et al. Assessing the quality of prediction models in healthcare using the Prediction Model Risk of Bias Assessment Tool (PROBAST): an evaluation of its use and practical application[J]. J Clin Epidemiol, 2025, 181: 111732. doi:10.1016/j.jclinepi.2025.111732

12. MoonsKGM, DamenJAA, KaulT, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods[J]. BMJ, 2025, 388: e082505. doi:10.1136/bmj-2024-082505

13. FinlaysonSG, BeamAL, van SmedenM. Machine learning and statistics in clinical research articles-moving past the false dichotomy[J]. JAMA Pediatr, 2023, 177(5): 448-450. doi:10.1001/jamapediatrics.2023.0034

14. HigginsJPT, ThomasJ, ChandlerJ, et al. Cochrane handbook for systematic reviews of interventions version 6.5 (updated August 2024)[M]. Cochrane: Cochrane Collaboration, 2024.

15. BurdickH, PinoE, Gabel-ComeauD, et al. Validation of a machine learning algorithm for early severe sepsis prediction: a retrospective study predicting severe sepsis up to 48 h in advance using a diverse dataset from 461 US hospitals[J]. BMC Med Inform Decis Mak, 2020, 20(1): 276. doi:10.1186/s12911-020-01284-x

16. 刘俊平. 诊断试验偏倚来源的研究进展[J]. 中国循证医学杂志, 2011, 11(7): 835-840.LiuJP. Advances in research on bias sources in diagnostic test[J]. Chinese Journal of Evidence-Based Medicine, 2011, 11(7): 835-840.

17. RileyRD, SnellKI, EnsorJ, et al. Minimum sample size for developing a multivariable prediction model: part II-binary and time-to-event outcomes[J]. Stat Med, 2019, 38(7): 1276-1296. doi:10.1002/sim.7992

18. RileyRD, SnellKIE, EnsorJ, et al. Minimum sample size for developing a multivariable prediction model: part I-continuous outcomes[J]. Stat Med, 2019, 38(7): 1262-1275. doi:10.1002/sim.7993

19. RileyRD, EnsorJ, SnellKIE, et al. Calculating the sample size required for developing a clinical prediction model[J]. BMJ, 2020, 368: m441. doi:10.1136/bmj.m441

20. MoonsKGM, WolffRF, RileyRD, et al. PROBAST: a tool to assess risk of bias and applicability of prediction model studies: explanation and elaboration[J]. Ann Intern Med, 2019, 170(1): W1-W33. doi:10.7326/m18-1377

21. MäenpääSM, KorjaM. Diagnostic test accuracy of externally validated convolutional neural network (CNN) artificial intelligence (AI) models for emergency head CT scans-a systematic review[J]. Int J Med Inform, 2024, 189: 105523. doi:10.1016/j.ijmedinf.2024.105523

22. ElahmediM, SawhneyR, GuadagnoE, et al. The state of artificial intelligence in pediatric surgery: a systematic review[J]. J Pediatr Surg, 2024, 59(5): 774-782. doi:10.1016/j.jpedsurg.2024.01.044

23. Carrasco-RibellesLA, Llanes-JuradoJ, Gallego-MollC, et al. Prediction models using artificial intelligence and longitudinal data from electronic health records: a systematic methodological review[J]. J Am Med Inform Assoc, 2023, 30(12): 2072-2082. doi:10.1093/jamia/ocad168

24. GanapathiS, PalmerJ, AldermanJE, et al. Tackling bias in AI health datasets through the STANDING together initiative[J]. Nat Med, 2022, 28(11): 2232-2233. doi:10.1038/s41591-022-01987-w

25. CollinsGS, ReitsmaJB, AltmanDG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement[J]. BMJ, 2015, 350: g7594. doi:10.1111/eci.12376

26. CollinsGS, MoonsKGM, DhimanP, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods[J]. BMJ, 2024, 385: e078378. doi:10.1136/bmj.q902

Popular Papers