A model nobody can run cannot be checked.
A published model is meant to be tested, questioned, and built on. That takes code that other researchers can run on their own data. In a review of 218 AI studies published in RSNA journals from 2017 through 2021, 34% shared code, and 11% shared code documented well enough to reproduce the study.1 An automated attempt to rerun 15,817 Jupyter notebooks linked to biomedical papers in PubMed Central reproduced the original results for 879.2
Testing on new data matters too. Of 86 deep learning algorithms for radiologic diagnosis that were evaluated on external data, 70 reported at least some decrease in performance there compared with internal data.3 We think a published model should be straightforward to run.
1. Venkatesh K, Santomartino SM, Sulam J, Yi PH. Code and data sharing practices in the radiology artificial intelligence literature: a meta-research study. Radiol Artif Intell. 2022;4(5):e220081.
2. Samuel S, Mietchen D. Computational reproducibility of Jupyter notebooks from biomedical publications. Gigascience. 2024;13:giad113.
3. Yu AC, Mohajer B, Eng J. External validation of deep learning algorithms for radiologic diagnosis: a systematic review. Radiol Artif Intell. 2022;4(3):e210064.
