Managing Reproducibility Risk in Multi-Omics Analytics

Authors

  • Niloofar Jamareh * Department of Biosciences, Università degli Studi di Milano, Via Celoria 26, 20133 Milan, Italy.
  • Neda Jamareh Department of Mathematics and Physics, University of Campania “Luigi Vanvitelli”, Viale Lincoln 5, 81100 Caserta, Italy.

https://doi.org/10.22105/masi.v3i4.115

Abstract

Multi-omics analytics can integrate genomic, transcriptomic, epigenomic, proteomic, metabolomic, and related data into richer representations of biological systems, but analytical flexibility can also make strong results sensitive to defensible changes in preprocessing, integration, validation, or context. This paper examines that problem as reproducibility risk at the decision level. We combine a structured narrative synthesis of multi-omics methodology, computational reproducibility, biomedical machine learning governance, and organizational decision research with an illustrative Monte Carlo stress test. We define the Analytical Fragility Index (AFI) as the proportion of decision outputs that change across pre-specified plausible perturbations. In 500 simulated experiments with three omics signal layers, increasing technical confounding made an internally selected pipeline look progressively stronger while becoming less transportable. At the highest confounding level, mean development Area Under The Receiver Operating Characteristic Curve (AUROC) reached 0.978, replication AUROC fell to 0.753, the optimism gap reached 0.225, and 29.8% of binary decisions changed after technical context shift. These are illustrative rather than empirical estimates. We therefore propose six risk-based evidence gates linking provenance, design alignment, validation, stress testing, independent challenge, and post-decision learning. The implication is that multi-omics programs should manage analytical quality by the stability and consequence of decisions, not by model accuracy alone.

Keywords:

Multi-omics analytics, Reproducibility risk, Data governance, Monte Carlo simulation, Analytical fragility, Decision management

References

  1. [1] Hasin, Y., Seldin, M., & Lusis, A. (2017). Multi-omics approaches to disease. Genome biology, 18, 83. https://doi.org/10.1186/s13059-017-1215-1

  2. [2] Karczewski, K. J., & Snyder, M. P. (2018). Integrative omics for health and disease. Nature reviews genetics, 19, 299-310. https://doi.org/10.1038/nrg.2018.4

  3. [3] Liu, F., Beck, S., Yang, L., Luo, H., & Zhang, K. (2026). Advancing AI for multi-omics and clinical data integration in basic and translational cancer research. Nature reviews cancer, 26, 497-512. https://doi.org/10.1038/s41568-026-00922-2

  4. [4] Cui, H., Tejada-Lapuerta, A., Brbić, M., Saez-Rodriguez, J., Cristea, S., Goodarzi, H., Lotfollahi, M., Theis, F. J., & Wang, B. (2025). Towards multimodal foundation models in molecular cell biology. Nature, 640, 623-633. https://doi.org/10.1038/s41586-025-08710-y

  5. [5] Ritchie, M. D., Holzinger, E. R., Li, R., Pendergrass, S. A., & Kim, D. (2015). Methods of integrating data to uncover genotype-phenotype interactions. Nature reviews genetics, 16, 85-97. https://doi.org/10.1038/nrg3868

  6. [6] Picard, M., Scott-Boyer, M. P., Bodein, A., Périn, O., & Droit, A. (2021). Integration strategies of multi-omics data for machine learning analysis. Computational and structural biotechnology journal, 19, 3735-3746. https://doi.org/10.1016/j.csbj.2021.06.030

  7. [7] Acharya, D., & Mukhopadhyay, A. (2024). A comprehensive review of machine learning techniques for multi-omics data integration: Challenges and applications in precision oncology. Briefings in functional genomics, 23(5), 549-560. https://doi.org/10.1093/bfgp/elae013

  8. [8] Brooks, T. G., Lahens, N. F., Mrčela, A., & Grant, G. R. (2024). Challenges and best practices in omics benchmarking. Nature reviews genetics, 25(5), 326-339. https://doi.org/10.1038/s41576-023-00679-6

  9. [9] Hu, Y., Wan, S., Luo, Y., Li, Y., Wu, T., Deng, W., ... & Qu, K. (2024). Benchmarking algorithms for single-cell multi-omics prediction and integration. Nature methods, 21(11), 2182-2194. https://doi.org/10.1038/s41592-024-02429-w

  10. [10] Fu, S., Wang, S., Si, D., Li, G., Gao, Y., & Liu, Q. (2025). Benchmarking single-cell multi-modal data integrations. Nature methods, 22, 2437-2448. https://doi.org/10.1038/s41592-025-02737-9

  11. [11] Rohart, F., Gautier, B., Singh, A., & Lê Cao, K. A. (2017). mixOmics: An R package for omics feature selection and multiple data integration. PLOS computational biology, 13(11), e1005752. https://doi.org/10.1371/journal.pcbi.1005752

  12. [12] Argelaguet, R., Velten, B., Arnol, D., Dietrich, S., Zenz, T., Marioni, J. C., Buettner, F., Huber, W., & Stegle, O. (2018). Multi-omics factor analysis - a framework for unsupervised integration of multi-omics data sets. Molecular systems biology, 14(6), e8124. https://doi.org/10.15252/msb.20178124

  13. [13] Argelaguet, R., Arnol, D., Bredikhin, D., Deloro, Y., Velten, B., Marioni, J. C., & Stegle, O. (2020). MOFA+: A statistical framework for comprehensive integration of multi-modal single-cell data. Genome biology, 21, 111. https://doi.org/10.1186/s13059-020-02015-1

  14. [14] Conesa, A., Madrigal, P., Tarazona, S., Gomez-Cabrero, D., Cervera, A., McPherson, A., ... & Mortazavi, A. (2016). A survey of best practices for RNA-seq data analysis. Genome biology, 17(1), 13. https://doi.org/10.1186/s13059-016-0881-8

  15. [15] Goh, W. W. B., Wang, W., & Wong, L. (2017). Why batch effects matter in omics data, and how to avoid them. Trends in biotechnology, 35(6), 498-507. https://doi.org/10.1016/j.tibtech.2017.02.012

  16. [16] Tung, P. Y., Blischak, J. D., Hsiao, C. J., Knowles, D. A., Burnett, J. E., Pritchard, J. K., & Gilad, Y. (2017). Batch effects and the effective design of single-cell gene expression studies. Scientific reports, 7, 39921. https://doi.org/10.1038/srep39921

  17. [17] Heumos, L., Schaar, A. C., Lance, C., Litinetskaya, A., Drost, F., Zappia, L., ... & Theis, F. J. (2023). Best practices for single-cell analysis across modalities. Nature reviews genetics, 24(8), 550-572. https://doi.org/10.1038/s41576-023-00586-w

  18. [18] Bernett, J., Blumenthal, D. B., Grimm, D. G., Haselbeck, F., Joeres, R., Kalinina, O. V., & List, M. (2024). Guiding questions to avoid data leakage in biological machine learning applications. Nature methods, 21(8), 1444-1453. https://doi.org/10.1038/s41592-024-02362-y

  19. [19] Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804

  20. [20] Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J. (2016). The FAIR guiding principles for scientific data management and stewardship. Scientific data, 3, 160018. https://doi.org/10.1038/sdata.2016.18

  21. [21] Sandve, G. K., Nekrutenko, A., Taylor, J., & Hovig, E. (2013). Ten simple rules for reproducible computational research. PLOS computational biology, 9(10), e1003285. https://doi.org/10.1371/journal.pcbi.1003285

  22. [22] Munafò, M. R., Nosek, B. A., Bishop, D. V., Button, K. S., Chambers, C. D., Percie du Sert, N., ... & Ioannidis, J. P. (2017). A manifesto for reproducible science. Nature human behaviour, 1(1), 0021. https://doi.org/10.1038/s41562-016-0021

  23. [23] Begley, C. G., & Ellis, L. M. (2012). Raise standards for preclinical cancer research. Nature, 483, 531-533. https://doi.org/10.1038/483531a

  24. [24] Freedman, L. P., Cockburn, I. M., & Simcoe, T. S. (2015). The economics of reproducibility in preclinical research. PLOS biology, 13(6), e1002165. https://doi.org/10.1371/journal.pbio.1002165

  25. [25] Collins, F. S., & Tabak, L. A. (2014). Policy: NIH plans to enhance reproducibility. Nature, 505, 612-613. https://doi.org/10.1038/505612a

  26. [26] McDermott, M. B. A., Wang, S., Marinsek, N., Ranganath, R., Foschini, L., & Ghassemi, M. (2021). Reproducibility in machine learning for health research: Still a ways to go. Science translational medicine, 13(586), eabb1655. https://doi.org/10.1126/scitranslmed.abb1655

  27. [27] Roberts, M., Driggs, D., Thorpe, M., Gilbey, J., Yeung, M., Ursprung, S., Aviles-Rivero, A. I., Etmann, C., McCague, C., Beer, L., Weir-McCall, J. R., Teng, Z., Gkrania-Klotsas, E., Aix-Covnet, Rudd, J. H. F., Sala, E., & Schönlieb, C. B. (2021). Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nature machine intelligence, 3, 199-217. https://doi.org/10.1038/s42256-021-00307-0

  28. [28] Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723

  29. [29] Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., ... & Gebru, T. (2019). Model cards for model reporting. Proceedings of the conference on fairness, accountability, and transparency (pp. 220-229). 2019 Association for Computing Machinery. https://doi.org/10.1145/3287560.3287596

  30. [30] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in neural information processing systems, 28, 2503-2511. https://proceedings.neurips.cc/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf

  31. [31] Futoma, J., Simons, M., Panch, T., Doshi-Velez, F., & Celi, L. A. (2020). The myth of generalisability in clinical research and machine learning in health care. The lancet digital health, 2(9), e489-e492. https://doi.org/10.1016/S2589-7500(20)30186-2

  32. [32] Finlayson, S. G., Subbaswamy, A., Singh, K., Bowers, J., Kupke, A., Zittrain, J., Kohane, I. S., & Saria, S. (2021). The clinician and dataset shift in artificial intelligence. New England journal of medicine, 385(3), 283-286. https://doi.org/10.1056/NEJMc2104626

  33. [33] Wiens, J., Saria, S., Sendak, M., Ghassemi, M., Liu, V. X., Doshi-Velez, F., Jung, K., Heller, K., Kale, D., Saeed, M., Ossorio, P. N., Thadaney-Israni, S., & Goldenberg, A. (2019). Do no harm: A roadmap for responsible machine learning for health care. Nature medicine, 25, 1337-1340. https://doi.org/10.1038/s41591-019-0548-6

  34. [34] Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G., & King, D. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC medicine, 17, 195. https://doi.org/10.1186/s12916-019-1426-2

  35. [35] Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., & Reitsma, J. B. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, e078378. https://doi.org/10.1136/bmj-2023-078378

  36. [36] Moons, K. G. M., Damen, J. A. A., Kaul, T., Hooft, L., Andaur Navarro, C., Dhiman, P. (2025). PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ, 388, e082505. https://doi.org/10.1136/bmj-2024-082505

  37. [37] Lekadir, K., Frangi, A. F., Porras, A. R., Glocker, B., Cintas, C., Langlotz, C. P. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ, 388, e081554. https://doi.org/10.1136/bmj-2024-081554

  38. [38] World Health Organization. (2021). Ethics and governance of artificial intelligence for health: WHO guidance. World Health Organization. https://iris.who.int/server/api/core/bitstreams/f780d926-4ae3-42ce-a6d6-e898a5562621/content?utm_medium=email&utm_source=transaction

  39. [39] Tabassi, E. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). https://doi.org/10.6028/NIST.AI.100-1

  40. [40] March, J. G. (1991). Exploration and exploitation in organizational learning. Organization science, 2(1), 71-87. https://doi.org/10.1287/orsc.2.1.71

  41. [41] Argote, L., & Miron-Spektor, E. (2011). Organizational learning: From experience to knowledge. Organization science, 22(5), 1123-1137. https://doi.org/10.1287/orsc.1100.0621

  42. [42] Berente, N., Gu, B., Recker, J., & Santhanam, R. (2021). Managing artificial intelligence. MIS quarterly, 45(3), 1433-1450. https://doi.org/10.25300/MISQ/2021/16274

  43. [43] Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California management review, 61(4), 66-83. https://doi.org/10.1177/0008125619862257

  44. [44] Faraj, S., Pachidi, S., & Sayegh, K. (2018). Working and organizing in the age of the learning algorithm. Information and organization, 28(1), 62-70. https://doi.org/10.1016/j.infoandorg.2018.02.005

Downloads

Published

2026-12-26

How to Cite

Jamareh, N. ., & Neda Jamareh. (2026). Managing Reproducibility Risk in Multi-Omics Analytics. Management Analytics and Social Insights, 3(4), 302-316. https://doi.org/10.22105/masi.v3i4.115

Similar Articles

1-10 of 66

You may also start an advanced similarity search for this article.