Prognostic Biomarker Discovery in Heart Failure
Predicting cardiovascular death among patients who already have heart failure, using multimodal machine learning on UK Biobank (N ≈ 500,000).

From risk to prognosis
Earlier Hurdle work modelled incident cardiovascular outcomes across a broad population and found cardiovascular death to be a particularly informative endpoint. The prognostic pipeline keeps that same CV-death endpoint, so results stay comparable, but shifts the question from "who will develop CV disease?" to "among people who already have heart failure, who will die of a cardiovascular cause?"
| Risk biomarkers | Prognostic biomarkers | |
|---|---|---|
| Population | Broad cohort by condition codes | Patients with I50.* heart failure |
| Cases | First ICD outcome after recruitment | Death from any CV cause (I00–I99) |
| Question | Who will get the disease? | Who will die from the disease? |
Risk biomarkers vs prognostic biomarkers — how the question changes.
Cohort and approach
- Inclusion: participants with an I50.* heart failure diagnosis (primary run requires HF at any time, with no diagnosis-based exclusions to preserve statistical power).
- Outcome: cases are HF patients whose cause of death is cardiovascular (ICD I00–I99); controls are HF patients who survived, were censored, or died of a non-CV cause.
- Model: penalised Cox time-to-event model across stacked data modalities, from routine clinical data through to proteomics and polygenic risk scores.
- Engineering: refactored per-period inclusion/exclusion with staged filtering and a can_use_in_cohort audit flag; stricter variants are one config edit away for sensitivity analysis.
Result 1: discrimination improves with each data layer

C-index for 15-year CV-death prognosis in HF patients, by data modality (UK Biobank).
Adding clinical labs and proteomics lifts the C-index from ~0.60 (baseline demographics) to 0.76 for the best multimodal model. Genetics alone adds little; the clinical + proteomic stack carries most of the gain, consistent with Hurdle's clinical-enrichment experiments on CV death. The cohort comprised 2,870 cardiovascular deaths and 13,506 sampled controls (~1:4.7); 73.2% of cases were male.
| Typology | C-index | ROC-AUC | Events / test n |
|---|---|---|---|
| Baseline + clinical + proteomics + PRS | 0.758 | 0.743 | 95 / 475 |
| Baseline + proteomics | 0.744 | 0.716 | 96 / 483 |
| Baseline + clinical | 0.672 | 0.657 | 563 / 3,260 |
| Baseline + standard PRS | 0.611 | 0.599 | 548 / 3,146 |
| Baseline (age, sex) | 0.599 | 0.588 | 568 / 3,276 |
Table 1 — C-index for 15-year CV-death prognosis in HF patients, by data modality (UK Biobank).
Result 2: multi-protein signatures beat a single marker

Per-SD hazard ratio for CV death, top vs bottom tertile (95% CI). UK Biobank HF cohort.
NT-proBNP is the canonical prognostic marker for heart failure, but on its own it gives only modest separation (HR 1.77 per SD). A multi-protein Cox signature roughly doubles the hazard ratio to 3.50 and widens 15-year survival separation between high- and low-risk tertiles from ~22 to ~46 percentage points. Prognosis is multi-dimensional; Hurdle's panel-level discovery captures signal that a single assay misses.
Discovered biomarkers tell a coherent clinical story
| Marker | Direction | Why it makes biological sense |
|---|---|---|
| NT-proBNP | ↑ risk | Canonical measure of cardiac dysfunction and congestion. |
| Renin | ↑ risk | RAAS activation; reduced renal perfusion worsens HF. |
| SPON1 | ↑ risk | ECM protein; marker of advanced fibrosis and remodelling. |
| ANGPTL4 / EDN1 | ↑ risk | Metabolic stress and vasoconstriction in failing hearts. |
| GDF-15 / urate | ↑ risk | Systemic stress and cardiorenal-metabolic burden. |
| HPGDS / diastolic BP | ↓ risk | Higher levels track with better prognosis. |
Table 2 — discovered biomarkers and their clinical rationale.
These are cardiovascular- and HF-specific signals (natriuretic peptides, RAAS, fibrosis and remodelling markers), not arbitrary noise. That coherence, together with C-index 0.76, makes the discovery both statistically useful and clinically interpretable.
Why it matters and where it goes next
The prognostic pipeline and per-period inclusion/exclusion filtering work together on a real UK Biobank cohort, turning the same Aurora engine that powers risk discovery into a prognostic-biomarker discovery platform. The framework generalises to any disease: a new cohort is defined by editing cohort JSON and enabling prognostic mode in the pipeline config, supporting trial enrichment, patient stratification, and companion-diagnostic development across therapeutic areas.

CEO, Hurdle
He/Him. Tom is CEO at Hurdle, a diagnostic-as-a-service company. Tom is a specialist in Epigenetics, Machine Learning, and Computational Biology.
LinkedIn →Contact us to discuss prognostic biomarker discovery and patient stratification for your programme.
Talk to us →