Real-world evaluation of large language models in healthcare (RWE-LLM): a new realm of AI safety & validation

Bhimani, M., Miller, A., Agnew, J.D., Ausin, M.S.

M Bhimani, A Miller, JD Agnew, MS Ausin… - medRxiv, 2025 - medrxiv.org

22 citations2025DOI: 10.1101/2025.03.17.25324157.abstract

Abstract

Background: The deployment of artificial intelligence (AI) in healthcare necessitates robust safety validation frameworks, particularly for systems directly interacting with patients. While theoretical frameworks exist, there remains a critical gap between abstract principles and practical implementation. Traditional LLM benchmarking approaches provide very limited output coverage and are insufficient for healthcare applications requiring high safety standards. Objective: To develop and evaluate a comprehensive framework for healthcare AI safety validation through large-scale clinician engagement.