Abstract
Autoregressive foundation models for electronic health records can make predictions for diverse clinical tasks without finetuning. These methods work by generating possible "synthetic futures" for patients, and then inferring downstream predictions over these simulated trajectories. Despite their strong capabilities, these methods have several flaws. (1) Inference is extremely computationally expensive, as each prediction requires simulating many trajectories, each with many observations. (2) Their performance is noisy due to the variance inherent to simulation, which in particular reduces efficacy on rare events, which are often of high importance clinically. (3) They are not promptable, as their only input is the patient's medical history, requiring indirect aggregation at inference time. We introduce EveryQuery, a promptable foundation model for structured EHR data. Building on the insight that clinical prediction tasks can be written in a structured query language, EveryQuery takes a query as a prompt alongside the patient's history and estimates the outcome directly. This allows EveryQuery to make predictions across diverse clinical tasks, such as classification and survival modeling, without simulation or finetuning. Across three EHR datasets, EveryQuery outperforms a competitive autoregressive baseline on clinically relevant classification and time-to-event tasks, with mean AUROC gains of 14.3% and 12.7% on MIMIC-IV and a large academic medical center and comparable performance on the smaller NWICU dataset. Its advantage is largest for rare outcomes, and it requires 580x less inference compute.