Garmin Metrics

Whose Recovery Score Should You Trust? Garmin, Whoop, Oura

August 1, 2026

Whose recovery score should you trust?

None of them more than the others, and not for the reason you expect. There is no published head-to-head validation of Garmin Body Battery against Whoop Recovery against Oura Readiness, and not one of the three companies publishes the formula that turns your heart rate data into the number on the screen.

That is not me being cute. In 2025 a research group at University College Dublin went looking for exactly this evidence. Doherty and colleagues catalogued 14 composite health scores across 10 wearable manufacturers — Body Battery and Training Readiness, Oura Readiness and Resilience, Whoop Recovery and Strain, plus Fitbit, Polar, Samsung, Suunto, Ultrahuman and Coros. Their finding on transparency was flat: no manufacturer disclosed an exact algorithmic formula, and few supplied peer-reviewed evidence that the score means anything.

So the honest comparison is not "which score is right." It is "what is each score built from, and how much independent information is left once you know that." On the second question I do have data, because our platform has deep Garmin histories and nothing equivalent from Whoop or Oura. I will be explicit every time that asymmetry matters.

What each score is actually computed from

Start with what the companies themselves say, because it is less than most people assume.

Garmin Body Battery. The owner's manual is one sentence: "Your watch analyzes your heart rate variability, stress level, sleep quality, and activity data to determine your overall Body Battery level." It runs 5 to 100, charges with sleep and drains with stress and effort, and updates all day rather than once at wake, which our full Body Battery explainer walks through screen by screen. Garmin's stress level is itself derived from HRV, which our Garmin stress score explainer covers, so two of the four listed inputs come off the same beat-to-beat signal. The underlying stress-and-recovery method comes from Firstbeat, whose white paper on 24-hour HRV analysis describes classifying the day into stress, recovery and physical activity periods from HRV alone.

Whoop Recovery. Whoop's developer documentation calls it "a daily measure of how prepared your body is to perform," calculated from "measurements from the previous day and your sleep, such as resting heart rate (RHR), heart rate variability (HRV), respiratory rate, sleep duration/quality, skin temperature, and blood oxygen." The Whoop 101 reference is the clearest primary source they publish. The measurement window is the night, not the day.

Oura Readiness. Oura publishes the most detail of the three: nine named contributors — resting heart rate, HRV balance, body temperature, recovery index, sleep, sleep balance, sleep regularity, previous day activity, and activity balance. Several are explicitly baseline-relative. HRV balance compares recent HRV against a longer trailing average; body temperature compares last night against your own long-term nighttime baseline. Oura describes the Readiness Score as combining overnight metrics with 14-day and two-month trends.

Line those up and the shared skeleton is obvious. Doherty's review quantified it across the whole category: HRV appeared in 86% of composite scores, resting heart rate in 79%, physical activity in 71%, sleep duration in 71%. Oura adds temperature. Whoop adds respiratory rate and SpO2. Garmin runs continuously instead of once a night. Underneath, all three are reading autonomic tone off a pulse signal and comparing it to your own history.

Our data: those inputs are close to one signal

Here is where I can put numbers on the overlap, on the Garmin side only.

We ran same-day within-athlete correlations across our platform's Garmin histories. Every figure below is a median of per-athlete Spearman rho, so it describes how the metrics move together inside one person rather than across a population.

  • Body Battery (day peak) vs overnight HRV: median rho 0.674, 82 athletes. Every single athlete came out positive, and 97.6% were above 0.3.
  • Body Battery vs Garmin sleep score: median rho 0.683, 85 athletes. Again 100% positive.
  • Control, HRV vs resting heart rate: median rho −0.696, 82 athletes. Textbook physiology, correct sign, which tells you the pipeline is measuring something real and not producing noise.
  • Control, Body Battery vs stress: median rho −0.452, 93 athletes.

The morning Body Battery peak is tracking your overnight HRV at roughly 0.67 and your sleep score at roughly 0.68, and HRV and resting heart rate are tracking each other at −0.70. That is not four opinions. It is one autonomic reading rendered four ways, with a few percent of independent variance sitting on top.

Cohort caveat, mandatory: these are self-selected Garmin owners who connected a training-analytics app. They are probably fitter, more consistent wearers, and more data-curious than the general population, and every number above should be read as "in this cohort."

And the asymmetry again, because it is the load-bearing limitation of this article: we have no Whoop or Oura data. I cannot tell you that Whoop Recovery correlates 0.67 with its own HRV input. What I can tell you is that Whoop and Oura draw on the same physiology, from the same autonomic system, measured over the same night, and there is no mechanism by which HRV and resting heart rate stop being two views of the same thing when you move the sensor to a finger. If you think Whoop or Oura escapes this, the burden is on the claim, not on the skeptic.

The composite did not predict performance

Correlated inputs would be forgivable if the output worked. In a separate study we checked Garmin's Training Readiness score, which sits downstream of Body Battery, HRV, sleep and recent load, against how athletes actually ran that day.

Across 5,136 paired steady runs from 39 athletes, the median within-athlete correlation between the morning readiness score and running efficiency was 0.056. Effectively zero. Splitting runs into readiness quintiles produced a spread of median efficiency z-scores from −0.12 to +0.07, which is nothing.

The positive control passed: readiness vs the previous day's training load came out at median rho −0.272, which it must, because recent load is an input to readiness by construction. So the null is a real null, not a broken join. We wrote up the practical side in why your Garmin training readiness score looks wrong.

My defensible opinion, stated plainly: a composite that is 0.67 correlated with one of its own inputs and 0.06 correlated with your actual output is not a readiness measurement. It is a nicely packaged restatement of last night's HRV. I think all three companies know this, which is why none of them publish a validation of the composite and all of them publish validations of the sensor instead.

Where the three genuinely differ: the sensor, not the formula

The one place independent head-to-head evidence exists is the raw signal, and it is worth reading carefully because it is the strongest argument any of these brands has.

Dial and colleagues (2025, Physiological Reports) had 13 adults sleep in five wearables at once against a Polar H10 single-lead ECG reference, across 536 nights. For overnight HRV (RMSSD) against ECG: Oura Gen 4 was best (CCC 0.99, MAPE 5.96%), then Oura Gen 3 (0.97, 7.15%), then Whoop 4.0 (0.94, 8.17%), then Garmin Fenix 6 (0.87, 10.52%), then Polar Grit X Pro (0.82, 16.32%). Garmin was excluded from the resting-heart-rate comparison because its timestamp method is undisclosed. The work was funded by the Air Force Research Laboratory, with no declared commercial conflicts.

Two honest readings of that. First, it is 13 people, and it compares specific model years, not brands forever. Second, site and device are confounded: Oura is a ring, Whoop and Garmin are wrist straps, and you cannot separate "finger is a better place to measure" from "Oura's algorithm is better" in this design. There is independent reason to think site matters — Charlton and colleagues showed in PLOS Digital Health that wrist PPG signal quality swings from 18.6 dB SNR lying down to 9.0 dB standing, and improves by roughly 5 dB just from raising the sensor to heart height. A finger, tightly coupled and well perfused, has an easier job. But "easier job" is a mechanism, not a measured brand verdict.

Sleep staging is the other place with real evidence, and it is again mostly Oura's. The Oura Gen3 OSSA 2.0 algorithm was tested against ambulatory polysomnography in 96 participants across 421,045 epochs (Sleep Medicine, 2024): 91.7–91.8% sleep/wake accuracy, per-stage accuracy from 75.5% for light sleep to 90.6% for REM. A formal comment was later published questioning parts of that analysis, and a response followed, so treat it as contested rather than settled.

So: measured evidence that Oura's ring reads HRV closer to ECG than a Fenix 6 does. Zero measured evidence that Oura Readiness tells you more about your day than Body Battery does. Those are different claims and the marketing for all three brands relies on you conflating them.

What to do instead

Pick one and stay on it. All three scores are baseline-relative. A device needs weeks of your own history before its number means anything, and switching resets that. If you are choosing between platforms for other reasons, our Garmin vs Apple Watch comparison for endurance athletes covers the ecosystem trade-offs.

Never treat two devices agreeing as confirmation. This is the practical point of everything above. If your ring and your watch both say 40 today, they have not independently verified each other. They have both read the same depressed overnight HRV and applied different arithmetic. Agreement between two instruments measuring the same underlying variable is expected, not evidential.

Read the input, not the composite. HRV status and resting heart rate are physiological measurements with decades of literature behind them. Body Battery is an unpublished transform of those measurements. When the two disagree, the input is the more defensible number. When Body Battery stops updating entirely, that is usually a data-capture problem rather than a physiological one.

Use it to catch outliers, not to plan. A score two standard deviations below your own baseline, alongside an elevated resting heart rate and a temperature bump, is worth acting on. A drop from 72 to 65 is noise wearing a number.


Frequently Asked Questions

Is Whoop Recovery more accurate than Garmin Body Battery?

Nobody has measured that. There is no published head-to-head validation of the two composite scores against any external standard. What has been measured is the underlying HRV signal: in Dial et al. (2025), Whoop 4.0 agreed with ECG better than a Garmin Fenix 6 (CCC 0.94 vs 0.87). That is evidence about the sensor, not about the score built on top of it.

Why do my Garmin and Oura scores disagree on the same morning?

Different measurement windows and different baselines. Body Battery updates continuously and reflects your day so far; Oura Readiness weights the previous night plus 14-day and two-month trends. Two algorithms with different memory lengths will disagree on any single day even when they see identical physiology. The disagreement tells you about the formulas, not about you.

Which recovery score should I buy a device for?

If HRV fidelity is your priority, the published evidence currently favours finger measurement, with the caveat that it comes from a 13-person study of specific device generations. If you already own a Garmin, the marginal information gain from adding a second device is smaller than the marketing suggests, because you would be buying a second look at the same signal.

Does a low recovery score mean I should skip my workout?

Not on its own. In our data the readiness composite had a median within-athlete correlation of 0.056 with same-day running efficiency across 5,136 runs. A low score plus how you actually feel plus an elevated resting heart rate is a reason to back off. A low score alone is one autonomic reading with a marketing layer on it.

Do any of these companies publish their algorithms?

No. Doherty et al. (2025) reviewed 14 composite scores across 10 manufacturers and found none disclosed an exact formula. Oura publishes the most, naming nine contributors without weights. Garmin names four inputs in one sentence. Whoop names six.


The interesting question is not which brand wins. It is why an industry that publishes ECG-validation studies for its sensors has published nothing at all for the scores those sensors feed. The sensor validation is the part that is easy to win. If someone releases a study showing their composite predicts anything a reader cares about — a race result, an illness, a bad session — that will be a genuine first, and we will cover it.

Keep Reading