Real-world evaluation of large language models in healthcare (RWE-LLM): a new realm of AI safety & validation
Bhimani, M., Miller, A., Agnew, J.D., Ausin, M.S.
M Bhimani, A Miller, JD Agnew, MS Ausin… - medRxiv, 2025 - medrxiv.org
22 citations2025DOI: 10.1101/2025.03.17.25324157.abstract
Abstract
Background: The deployment of artificial intelligence (AI) in healthcare necessitates robust safety validation frameworks, particularly for systems directly interacting with patients. While theoretical frameworks exist, there remains a critical gap between abstract principles and practical implementation. Traditional LLM benchmarking approaches provide very limited output coverage and are insufficient for healthcare applications requiring high safety standards. Objective: To develop and evaluate a comprehensive framework for healthcare AI safety validation through large-scale clinician engagement.