Study S2 · August 2026

Does Training Readiness predict how a run goes?

Key finding

Across 5,136 steady runs, the median within-athlete correlation between Garmin's Training Readiness score and actual running efficiency was 0.056. A control built into the same pipeline came back clearly negative, so the null is a real result rather than a broken method.

Gneta, August 2026 · our athletes · 5,136 paired runs

39

athletes

5,136

paired steady runs

0.056

median correlation

What we did

We joined each athlete's morning Training Readiness score to the runs they actually did that day, keeping only steady efforts so that a hard interval session could not masquerade as a bad day. That left 5,136 paired runs across the cohort with at least 20 pairs each.

“Efficiency” here is speed divided by average heart rate on steady runs — 20 to 90 minutes, average HR between 120 and 175, over 3 km — then z-scored within each athlete so people are only ever compared against themselves. It is a rough proxy for how good a run felt in physiological terms, not a lab measurement.

For each athlete separately we computed the Spearman rank correlation between readiness and efficiency, then looked at the distribution of those correlations. Nobody is pooled with anybody else.

What we found

Finding 1

The correlation is indistinguishable from noise

Median within-athlete correlation: 0.056, with the middle half of athletes falling between −0.101 and 0.161. Only 59% of athletes were even positive — close to the 50% you would get from a coin.

Just 12.8% of athletes cleared a correlation of 0.3, the rough threshold where a relationship starts being useful for a decision. Not one athlete came in below −0.3.

Finding 2

Sorting runs by readiness barely moves the outcome

We split each athlete's runs into readiness quintiles and looked at efficiency, z-scored within that athlete. Bottom quintile median: −0.12. Top quintile: +0.03.

That is about a seventh of a standard deviation across the entire range of the score — from the days Garmin says you are least ready to the days it says you are most. The middle quintiles do not even order cleanly.

Finding 3

The control worked, which is what makes the null informative

Training Readiness includes recent training load by construction, so readiness against yesterday's load should come out clearly negative. It did: median −0.272 across the cohort, with only 3.7% of athletes positive.

This matters more than it looks. A null result is worthless if your pipeline simply cannot detect a relationship. Ours detected the one that had to be there, in the same data, with the same code — and still found nothing against performance.

Limits you should know before citing this

  • The cohort is self-selected. Garmin owners who sought out a third-party analytics tool. Fitter and more data-curious than average, and not a random sample of anyone.
  • Efficiency is a proxy, not a verdict on the run. Speed per heartbeat is confounded by terrain, heat, wind, surface and how hard the athlete chose to push. A score could be genuinely useful for something this measure cannot see.
  • Readiness may change behaviour. An athlete who sees a low score and runs easier partly erases the relationship we are trying to measure. This cuts against finding a correlation and we cannot separate it out.
  • Steady runs only. Filtering to 20–90 minute efforts in a moderate heart-rate band removes races and hard intervals — exactly the sessions where readiness might matter most.
  • Six accounts excluded. Near-identical multi-year histories under different Garmin accounts, created inside a known incident window. See the methodology.
  • Observational. No causal claim is made anywhere in this study.

The data

The aggregate output and the script that produced it are public. Every number on this page traces to that file.

View the aggregate data (JSON) →