Search | Search by Center | Search by Source | Keywords in Title
Zalikha AK, Hong TS, Small EA, Constant M, Harris AHS, Giori NJ. Can a Large Language Model Interpret Data in the Electronic Health Record to Infer Minimum Clinically Important Difference Achievement of Knee Osteoarthritis Outcome Score-Joint Replacement Score Following Total Knee Arthroplasty? The Journal of arthroplasty. 2025 Jul 1; 40(7S1):S153-S157, DOI: 10.1016/j.arth.2025.03.049.
Dimensions for VA is a web-based tool available to VA staff that enables detailed searches of published research and research projects. BACKGROUND: Obtaining total knee arthroplasty patient-reported outcomes for quality assessment is costly and difficult. We asked whether a large language model (LLM) could interpret electronic health record notes to differentiate patients attaining a 1-year minimum clinically important difference (MCID) for the Knee Osteoarthritis Outcome Score-Joint Replacement (KOOS-JR) from those who did not. We also investigated whether sufficient information to infer MCID achievement exists in the chart by having a blinded orthopaedic surgeon make the same determination. METHODS: In this retrospective case-control study, we selected 40 total knee arthroplasty patients who achieved 1-year KOOS-JR MCID and 40 who did not. Orthopaedic, emergency medicine, and primary care notes from zero to six months preoperatively and nine to 15 months postoperatively were deidentified. ChatGPT 3.5 (ChatGPT) interpreted these notes to determine whether the patient improved after surgery. A blinded orthopaedic surgeon classified these patients using all chart information. The sensitivity, specificity, and accuracy of ChatGPT and the surgeon''s responses were calculated. RESULTS: ChatGPT classified 78 of 80 cases with 97% sensitivity, but only 33% specificity. The surgeon''s assessment had 90% sensitivity and 63% specificity. Given the equal distribution of patients meeting or not meeting MCID, Chat GPT''s accuracy was 65%. The surgeon''s was 76%. CONCLUSIONS: ChatGPT''s assessment of KOOS-JR MCID attainment had 97% sensitivity, but only 33% specificity. False positives were commonly due to the LLM not having access to, or not properly interpreting, signs of problems in the chart. This was an initial evaluation of the current ability of a general-purpose LLM to evaluate patient outcomes based on information in chart notes. An orthopaedic surgeon''s assessment of the full chart suggests an opportunity to improve on this baseline performance, possibly enabling quality monitoring and identification of best practices across a large health care system. Additional work is needed to optimize model performance and confirm the utility of this approach.