Archive \ Volume.17 2026 Issue 1

Evidence-Bound Generation for Medicines Information with Verifiable Claims and Visible Uncertainty

, , ,
  1. Department of Medicines Information and Evidence Generation, Faculty of Pharmacy, University of Bologna, Bologna, Italy.
  2. Department of Verifiable AI Claims in Pharmacy, Faculty of Pharmacy, University of Turin, Turin, Italy.
  3. Department of Uncertainty Communication in Pharmacy AI, Faculty of Pharmacy, Sapienza University of Rome, Rome, Italy.

Abstract

Generative artificial intelligence can produce fluent medicines information while obscuring whether individual claims are supported, qualified, contradicted, or unsupported by current evidence. Retrieval augmentation may improve access to external information, but retrieval alone does not establish source eligibility, claim faithfulness, clinical applicability, or safe decision use. This article develops an original, non-empirical Retrieval–Evidence–Generation Architecture for evidence-bound medicines information. Evidence-bound generation is defined as a proposed form of generation in which externally verifiable claims are constrained by eligible evidence, linked to inspectable evidence units, assigned explicit support states, and accompanied by visible uncertainty, conflict, or evidence absence. The architecture separates task specification, governed retrieval, evidence construction, claim planning, bounded generation, claim-level verification, professional review, and lifecycle monitoring. It further proposes a provenance chain connecting each displayed claim to its supporting passage, source identity, version, retrieval context, and transformation history. Five qualitative claim states—supported, supported with qualification, conflicting evidence, unsupported, and evidence absent—are proposed to prevent citations or model confidence from functioning as substitutes for verification. Validation would require technical assessment of retrieval and evidence binding, pharmacist assessment of medication-content correctness and harmful omission, human-factors testing of reliance and verification burden, equity assessment, and prospective workflow evaluation. The architecture does not establish clinical benefit, medication-safety improvement, regulatory acceptance, or deployment readiness. Its original contribution is to make evidence control, uncertainty, professional authority, and governance integral to medicines-information generation rather than optional post-generation checks.


Downloads: 20
Views: 92

How to cite:
Vancouver
Ferraro L, Ricci M, Moretti G, Greco P. Evidence-Bound Generation for Medicines Information with Verifiable Claims and Visible Uncertainty. Arch Pharm Pract. 2026;17(1):48-56. https://doi.org/10.51847/TzCg1KSqsZ
APA
Ferraro, L., Ricci, M., Moretti, G., & Greco, P. (2026). Evidence-Bound Generation for Medicines Information with Verifiable Claims and Visible Uncertainty. Archives of Pharmacy Practice, 17(1), 48-56. https://doi.org/10.51847/TzCg1KSqsZ

Download Citation
References
  1. Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233-9. doi:10.1056/NEJMsr2214184
  2. Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172-80. doi:10.1038/s41586-023-06291-2
  3. Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-96. doi:10.1001/jamainternmed.2023.1838
  4. Morath B, Chiriac U, Jaszkowski E, Deiß C, Nürnberg H, Hörth K, et al. Performance and risks of ChatGPT used in drug information: An exploratory real-world analysis. Eur J Hosp Pharm. 2024;31(6):491-7. doi:10.1136/ejhpharm-2023-003750
  5. Huang J, Yang DM, Rong R, Nezafati K, Treager C, Chi Z, et al. A critical assessment of using ChatGPT for extracting structured data from clinical notes. NPJ Digit Med. 2024;7(1):106. doi:10.1038/s41746-024-01079-8
  6. Sutton RT, Pincock D, Baumgart DC, Sadowski DC, Fedorak RN, Kroeker KI. An overview of clinical decision support systems: Benefits, risks, and strategies for success. NPJ Digit Med. 2020;3:17. doi:10.1038/s41746-020-0221-y
  7. Wiens J, Saria S, Sendak M, Ghassemi M, Liu VX, Doshi-Velez F, et al. Do no harm: A roadmap for responsible machine learning for health care. Nat Med. 2019;25(9):1337-40. doi:10.1038/s41591-019-0548-6
  8. Amann J, Blasimme A, Vayena E, Frey D, Madai VI. Explainability for artificial intelligence in healthcare: A multidisciplinary perspective. BMC Med Inform Decis Mak. 2020;20(1):310. doi:10.1186/s12911-020-01332-6
  9. Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28(5):924-33. doi:10.1038/s41591-022-01772-9
  10. Grossman S, Zerilli T, Nathan JP. Appropriateness of ChatGPT as a resource for medication-related questions. Br J Clin Pharmacol. 2024;90(10):2691-5. doi:10.1111/bcp.16212
  11. Bhayana R, Fawzy A, Deng Y, Bleakney RR, Krishna S. Retrieval-augmented generation for large language models in radiology: Another leap forward in board examination performance. Radiology. 2024;313(1). doi:10.1148/radiol.241489
  12. Zhang G, Xu Z, Jin Q, Chen F, Fang Y, Liu Y, et al. Leveraging long context in retrieval augmented language models for medical question answering. NPJ Digit Med. 2025;8(1):239. doi:10.1038/s41746-025-01651-w
  13. Das S, Ge Y, Guo Y, Rajwal S, Hairston J, Powell J, et al. Two-layer retrieval-augmented generation framework for low-resource medical question answering using Reddit data: Proof-of-concept study. J Med Internet Res. 2025;27. doi:10.2196/66220
  14. Jin Q, Kim W, Chen Q, Comeau DC, Yeganova L, Wilbur WJ, et al. MedCPT: Contrastive pre-trained transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval. Bioinformatics. 2023;39(11). doi:10.1093/bioinformatics/btad651
  15. Tayebi Arasteh S, Lotfinia M, Bressem K, Siepmann R, Adams LC, Ferber D, et al. RadioRAG: Online retrieval-augmented generation for radiology question answering. Radiol Artif Intell. 2025;7(4). doi:10.1148/ryai.240476
  16. Chelli M, Descamps J, Lavoué V, Trojani C, Azar M, Deckert M, et al. Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. J Med Internet Res. 2024;26. doi:10.2196/53164
  17. Gorenshtein A, Shihada K, Sorka M, Aran D, Shelly S. LITERAS: Biomedical literature review and citation retrieval agents. Comput Biol Med. 2025;192(Pt B):110363. doi:10.1016/j.compbiomed.2025.110363
  18. Marshall IJ, Nye B, Kuiper J, Noel-Storr A, Marshall R, Maclean R, et al. Trialstreamer: A living, automatically updated database of clinical trial reports. J Am Med Inform Assoc. 2020;27(12):1903-12. doi:10.1093/jamia/ocaa163
  19. Munafò MR, Nosek BA, Bishop DVM, Button KS, Chambers CD, Percie du Sert N, et al. A manifesto for reproducible science. Nat Hum Behav. 2017;1(1):0021. doi:10.1038/s41562-016-0021
  20. Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. 2021;3(11). doi:10.1016/S2589-7500(21)00208-9
  21. Savage T, Wang J, Gallo R, Boukil A, Patel V, Safavi-Naini SAA, et al. Large language model uncertainty proxies: Discrimination and calibration for medical diagnosis and treatment. J Am Med Inform Assoc. 2025;32(1):139-49. doi:10.1093/jamia/ocae254
  22. Kompa B, Snoek J, Beam AL. Second opinion needed: Communicating uncertainty in medical machine learning. NPJ Digit Med. 2021;4(1):4. doi:10.1038/s41746-020-00367-3
  23. Begoli E, Bhattacharya T, Kusnezov D. The need for uncertainty quantification in machine-assisted medical decision making. Nat Mach Intell. 2019;1(1):20-3. doi:10.1038/s42256-018-0004-1
  24. Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW; Topic Group “Evaluating diagnostic tests and prediction models” of the STRATOS initiative. Calibration: The Achilles heel of predictive analytics. BMC Med. 2019;17(1):230. doi:10.1186/s12916-019-1466-7
  25. Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and evaluation of health care applications of large language models: A systematic review. JAMA. 2025;333(4):319-28. doi:10.1001/jama.2024.21700
  26. Norgeot B, Quer G, Beaulieu-Jones BK, Torkamani A, Dias R, Gianfrancesco M, et al. Minimum information about clinical artificial intelligence modeling: The MI-CLAIM checklist. Nat Med. 2020;26(9):1320-4. doi:10.1038/s41591-020-1041-y
  27. Miao BY, Chen IY, Williams CYK, Davidson J, Garcia-Agundez A, Sun S, et al. The MI-CLAIM-GEN checklist for generative artificial intelligence in health. Nat Med. 2025;31(5):1394-8. doi:10.1038/s41591-024-03470-0
  28. Nagendran M, Chen Y, Lovejoy CA, Gordon AC, Komorowski M, Harvey H, et al. Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. BMJ. 2020;368. doi:10.1136/bmj.m689
  29. Reddy S, Allan S, Coghlan S, Cooper P. A governance model for the application of AI in health care. J Am Med Inform Assoc. 2020;27(3):491-7. doi:10.1093/jamia/ocz192
  30. Reddy S. Generative AI in healthcare: An implementation science informed translational path on application, integration and governance. Implement Sci. 2024;19(1):27. doi:10.1186/s13012-024-01357-9
  31. Greenhalgh T, Wherton J, Papoutsi C, Lynch J, Hughes G, A’Court C, et al. Beyond adoption: A new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies. J Med Internet Res. 2017;19(11). doi:10.2196/jmir.8775
  32. Damschroder LJ, Reardon CM, Widerquist MAO, Lowery J. The updated Consolidated Framework for Implementation Research based on user feedback. Implement Sci. 2022;17(1):75. doi:10.1186/s13012-022-01245-0
  33. Lyell D, Coiera E. Automation bias and verification complexity: A systematic review. J Am Med Inform Assoc. 2017;24(2):423-31. doi:10.1093/jamia/ocw105
  34. Subbaswamy A, Saria S. From development to deployment: Dataset shift, causality, and shift-stable models in health AI. Biostatistics. 2020;21(2):345-52. doi:10.1093/biostatistics/kxz041
  35. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-53. doi:10.1126/science.aax2342
  36. Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019;17(1):195. doi:10.1186/s12916-019-1426-2
 

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.