Pulse.

a daily field guide to health research that matters

◆ Console

‹ Fri · 14 Aug 2026
Near-term implementable finding

Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States

Machine learning extracted cardiovascular events from medical records more accurately than traditional diagnostic codes, improving future event prediction in cancer patients.

A multisite retrospective validation study at Mayo Clinic compared zero-shot LLM extraction against ICD-code-based retrieval for identifying cardiovascular events in 3,684 patients across two independent cohorts. LLM achieved highest AUC for stroke (0.920), MI (0.938), and composite MACE (0.880) in the ICI-treated cohort, outperforming ICD coding for stroke and MACE identification in both cohorts.

What the study was

Study design
multisite_retrospective_validation
Population
ICI-treated patients (Cohort 1) and TAVR patients (Cohort 2) at Mayo Clinic
Sample size
3684
Category
Diagnostics
Maturity
Validated
Journal
BMJ Open

Why it surfaced

BMJ Open, CC BY-NC; multisite Mayo Clinic validation (n=3,684 total, 2 independent cohorts); benchmark study comparing LLM vs ICD code extraction; high AUCs for stroke/MACE with manual adjudication as gold standard; directly applicable to real-world clinical informatics workflows

A plain-language summary of published research — not medical advice. Talk to a clinician about your own care.