Garmin Metrics

Does Garmin Training Readiness Predict Anything? We Checked

August 1, 2026

Does Garmin's Training Readiness score actually predict anything?

It predicts your recent training load accurately and your next run barely at all. Across 5,136 steady runs from 39 athletes on our platform, the median within-athlete correlation between morning Training Readiness and how efficiently that day's run actually went was 0.056 — close enough to zero that you should not plan a Tuesday around it.

A null result is only worth reading if the method can find a signal when one exists. Ours can, and proving it is the whole reason I trust this finding, so that check gets its own section below.

This article is about whether the score forecasts performance. If your score is stuck at 1 and you want it fixed, that is a different problem, covered in why Training Readiness is always low and six fixes that work. For what the score is built from, start with our Training Readiness explainer.

What we measured

Training Readiness makes an implicit promise: a high number means today is a good day to go hard, a low number means it is not. If that holds, better-readiness days should produce measurably better runs. We tested exactly that on our own platform's data, in a study run on 2026-08-01:

  • The runs. 5,136 steady runs: 20 to 90 minutes, average heart rate between 120 and 175 bpm, longer than 3 km. The filter deliberately throws out intervals, races, and 12-minute shakeouts, because those tell you about the session design, not the athlete.
  • The athletes. 39, each with at least 20 paired run-and-readiness days. Six accounts were excluded outright: created during a 2026-05-03 incident window with near-identical multi-year histories under distinct Garmin accounts, so we treated them as non-independent rather than let them inflate an n.
  • The outcome. Running efficiency, defined as speed divided by average heart rate. More speed per heartbeat is a better run. Every value was z-scored within each athlete, so nobody is compared to anybody else.
  • The pairing. Each run matched to the Training Readiness score from that same day.

The within-athlete z-scoring matters more than it sounds. A fitter runner always has better raw efficiency, which tells you nothing about whether their watch predicted anything. The only question here is whether you run better than your average on your higher-readiness days.

The result: essentially flat

Median within-athlete Spearman correlation between readiness and efficiency: 0.056. For the shape of the spread: 59% of athletes came out positive, only 12.8% cleared a correlation of 0.3, and not a single athlete in the cohort came out below -0.3. Nobody's watch was backwards. It was just quiet.

Splitting each athlete's own days into readiness quintiles gives the same picture in a form you can actually read:

Readiness quintile (each athlete's own) Median efficiency z-score Runs
Q1 (lowest readiness) -0.12 1,020
Q2 +0.015 1,007
Q3 +0.065 1,017
Q4 +0.067 986
Q5 (highest readiness) +0.026 986

Two things stand out. The total spread from worst to best quintile is under 0.2 standard deviations. And the top quintile was no better than the middle: Q5 came in at +0.026, below both Q3 and Q4. The only quintile that looks different is the bottom one, and even that is a 0.12 SD dip.

If the score worked the way the interface implies, this column would climb. It does not. It steps up slightly out of the basement and then goes flat.

The positive control, and why it is the most important number here

Fail to find a signal with a broken ruler and you have measured your ruler, not the world. So before publishing "readiness doesn't predict performance," we had to show the same pipeline finding a relationship we already know is there.

Training Readiness includes recent training load by construction. It is one of the inputs listed in our breakdown of the score, so a hard session yesterday must push today's score down. That is not a hypothesis about physiology, it is arithmetic inside the algorithm. So if our method is sound, correlating readiness against yesterday's training duration has to come out clearly negative.

It did. Median within-athlete correlation: -0.272, and 96.3% of athletes came out negative (54 athletes had enough data for this particular check). Same code, same z-scoring, same cohort, and a clean correctly-signed relationship where one is guaranteed to exist.

That is what turns our flat performance result from "we found nothing" into "there is nothing there to find at this resolution." The method works. It just has nothing to report about tomorrow's run.

I would like more studies of consumer wearable metrics to include a control like this, our own earlier ones included. It costs one extra query and it is the difference between a finding and a shrug.

The strangest number: athletes train hardest on their worst days

One more relationship fell out of the data, and it points at people rather than at the watch. Readiness versus the same day's aerobic Training Effect came out at a median correlation of -0.209, with 43.6% of athletes below -0.3. The lower the readiness score, the harder the session tended to be.

The tempting interpretation is that the watch is inverted. It is not. This is behaviour, with two ordinary explanations:

  1. Hard blocks are multi-day. Yesterday's session pushed today's readiness down, and today you are on day two of a three-day block, so you go hard again. Readiness reports the consequence of Monday while you execute Tuesday. Both numbers are right; they describe different things.
  2. People ignore the watch. The plan says intervals, so the athlete does intervals. A number on a wrist rarely beats a schedule, a training group, or the one weekday evening that works.

Either way, the correlation tells you how athletes behave around the score, not whether the score is valid. Our training load calculator shows how acute and chronic load stack up across a block, which is where the multi-day pattern becomes obvious in your own data.

What this test cannot see

Four caveats, and they are real ones.

Single-day running efficiency is a blunt outcome. Speed per heartbeat on one run is confounded by terrain, heat, wind, surface, and intent. A run on a 25°C afternoon into a headwind scores badly no matter how recovered you are. A genuine but small readiness effect could easily hide under that noise. 0.056 is not proof of zero; it is proof that any effect is smaller than the everyday variance you already live with.

This is about acute daily prediction only. We tested whether today's score forecasts today's run. We did not test whether chronic recovery tracking has value over months, whether readiness catches illness early, or whether the trend spots overreaching before you feel it. This study says nothing about those.

The cohort is self-selected. These are Garmin owners who connected a training-analytics app: almost certainly fitter and more data-curious than the average watch wearer. Their wear-time discipline is probably better too, which if anything should have made the score more predictive here, not less.

Thirty-nine athletes is not a population. Enough to say the effect is not large and not consistent. Not enough to rule out a small effect in some subgroup we cannot see.

What to do instead

Stop using the number as a verdict and start using it as an accounting readout. It is genuinely good at what it actually does.

Read it backwards, not forwards. A low score is reliable evidence that you did something demanding recently, or slept badly, or both. That is confirmed by our positive control at -0.272. It is a receipt for the last few days, and receipts are useful.

Do not cancel a planned session on the score alone. Our data cannot support that decision. Top-quintile readiness days ran 0.15 SD better than bottom-quintile days, and slightly worse than middle-quintile days. Skip every session under 40 and you will skip a lot of perfectly good running.

Do let it break ties. If the session is already ambiguous, a 32 is a reasonable thumb on the scale toward the easier option. That is a different claim from "the score knows."

Watch the trend, not the morning. Same lesson we landed on measuring the race predictor: the direction over weeks beats any single day's number. A readiness score that has drifted from a 70s baseline to the 40s over ten days tells you something a single 44 never could. Pair it with Training Status for the load side of the same story.

If one input looks wrong, fix the input. Readiness is downstream of HRV, sleep, and recovery time. If your HRV status has been unbalanced for weeks or recovery time keeps coming back absurdly high, the score is faithfully reporting a broken input.

The honest verdict

Training Readiness is internally consistent and it reflects your recent load. It just does not forecast how a given run will go. Useful as a load-accounting readout; not useful as a daily go/no-go oracle.

My opinion, stated plainly: the presentation is what is wrong here, not the arithmetic. Labels like "Prime" and "Ready" promise a forecast that the number underneath cannot make, and athletes reasonably take the promise at face value. Call it what it is, a recent-load and sleep summary, and it becomes a trustworthy instrument instead of an oracle that quietly fails most times you consult it.


Frequently Asked Questions

Is Garmin Training Readiness accurate?

Accurate at what? It accurately reflects your recent training load, sleep, and HRV, which are its inputs; our positive control confirmed the load relationship at a median correlation of -0.272 across 54 athletes. It was not accurate at predicting performance: median correlation with same-day running efficiency was 0.056 across 5,136 runs.

Should I skip a workout when Training Readiness is low?

Not on the score alone. In our data, runs on athletes' lowest-readiness days were only about 0.12 standard deviations below their own average, and top-quintile days were no better than middle-quintile days. Use the score to break a tie when you already feel unsure, not to overrule a plan you were otherwise ready to execute.

Why is my Training Readiness low when I feel fine?

Usually because it is reporting your last 24-72 hours rather than your current state, and the most common single cause is a data-quality problem rather than physiology. Our troubleshooting guide covers the six fixes for a persistently low score, starting with sleep consistency and HRV data quality.

Is Training Readiness better than Body Battery for deciding when to train?

They share inputs, so they rarely disagree by much. We have not run the equivalent accuracy study on Body Battery yet, so we cannot tell you which predicts better, and we are not going to guess. If yours is behaving strangely, Body Battery not working and our Body Battery guide cover the mechanics.

Why do I do my hardest sessions on my lowest-readiness days?

Because hard training comes in multi-day blocks, so yesterday's session depresses today's score while today's session is still on the plan. We saw this across the cohort: readiness versus same-day aerobic Training Effect ran at a median correlation of -0.209, with 43.6% of athletes below -0.3.


There is a version of this metric that would earn the "should I train hard today" framing, and it needs something Garmin does not have: your own history of what happened after each score, weighted to you. Thirty-nine athletes told us the population-level answer is flat. That leaves the possibility that it is not flat for you, settleable only by keeping your own paired record for a few months and looking. Less satisfying than a number on a wrist. Also the true one.

Related reading:

Keep Reading