Garmin Metrics
Whose Sleep Score Should You Trust? Garmin, Whoop, Oura, Apple
August 2, 2026
The short answer
Apple's, and only because Apple publishes the arithmetic. The Apple Watch sleep score is 50 points for how long you slept, 30 for bedtime consistency against your last 13 nights, and 20 for interruptions, stated plainly in Apple's own support documentation. You can reconstruct it from your own data and check whether the number agrees with you.
None of the other three can be reconstructed. Garmin, Whoop and Oura all fold sleep-stage estimates and heart-rate-derived recovery signals into a single 0–100 figure, and none of them publishes the weights. That matters more than it sounds, because the stage estimates are the weakest measurement in the whole stack. In the best recent head-to-head against clinical polysomnography, agreement on four-stage sleep classification ranged from κ = 0.53 down to κ = 0.21 across six devices.
A note on what we can and cannot claim. Our own study data covers Garmin only, because Garmin is the only platform we hold deep histories for. Everything below about Whoop, Oura and Apple comes from those companies' documentation and from published validation work, linked where it appears. Our depth on Garmin exceeds our depth on the other three by a wide margin, and I will say so again at the point where it matters most.
What each score is actually built from
Garmin
Garmin's description of the sleep score is unusually specific about the ingredients and silent about the recipe. Garmin says the score is "calculated based on a blend of how long you slept, how well you slept and evidence of recovery activity occurring in your autonomic nervous system derived from heart rate variability data," and that quantity and quality are graded "by comparing your recorded sleep to age-based standards agreed upon by sleep experts" (Garmin's sleep score and sleep insights page).
The stages themselves come from "a combination of heart rate, heart rate variability, respiration rate, body movement and other key inputs." The third component, overnight recovery, comes off the same parasympathetic-versus-sympathetic analysis that drives all-day stress and Body Battery, which means your sleep score and your HRV status are partly reading the same beat-to-beat signal.
Whoop
Whoop's developer docs confirm the sleep stages it reports (Light, REM, Slow Wave Sleep) and that sleep need is computed "based on your Sleep Debt and your previous day's activity" (Whoop 101). Whoop rebuilt Sleep Performance in 2025 from a single hours-versus-needed ratio into a composite of sleep sufficiency, sleep consistency, sleep efficiency and overnight sleep stress.
Oura
Oura documents more of its structure than anyone except Apple. The Sleep Score has seven named contributors: total sleep, efficiency, restfulness, REM sleep, deep sleep, latency and timing. Oura says the underlying signals are "resting heart rate, average body temperature, movement, and time spent in specific sleep stages," and that total sleep time is the largest contributor, without saying how much larger.
Four of those seven contributors depend on the sleep-staging algorithm being right. That is a lot of weight on the part of the measurement with the worst published agreement.
Apple
Apple is the outlier, and the interesting thing is what it left out. Apple Watch estimates sleep stages, and Apple publishes a technical white paper on how — 1,171 nights from 858 participants to develop the classifier, 299 nights from 166 held-out participants to validate it, against in-lab PSG, at-home PSG and at-home EEG. Then Apple declined to put any of it in the score. No REM target, no deep sleep percentage, no HRV term. Three inputs, published weights, arithmetic you can do yourself.
What the validation literature actually shows
Two studies are worth your time here.
Chinoy and colleagues, in SLEEP (2021), put seven consumer devices next to polysomnography for 34 healthy adults across three consecutive lab nights including a disrupted-sleep condition. Every device detected sleep well: sensitivity ≥ 0.93 across the board. Every device was bad at detecting wake, with specificity from 0.54 (Fitbit Alta HR) down to 0.18 (Garmin Fenix 5S) and 0.19 (Garmin Vivosmart 3). On stages, the authors wrote that most devices "failed to correctly identify 30%–50% of both deep sleep and REM sleep, on average."
Schyvens and colleagues, in Sleep Advances (2025), ran six current-generation wrist devices against PSG in 62 adults (52 men, mean age 46.0 ± 12.6) at a sleep centre, a mix of healthy participants and people with suspected sleep apnea. Four-stage agreement, as Cohen's κ:
| Device | Cohen's κ vs PSG |
|---|---|
| Apple Watch Series 8 | 0.53 |
| Fitbit Sense | 0.42 |
| Fitbit Charge 5 | 0.41 |
| Whoop 4.0 | 0.37 |
| Withings Scanwatch | 0.22 |
| Garmin Vivosmart 4 | 0.21 |
Whoop 4.0 was the best at deep sleep, correctly classifying 69.63% of PSG N3 epochs. Apple was the best at REM at 68.57%. Both got the totals wrong in opposite directions: Apple underestimated deep sleep by 25.20 minutes a night, Whoop overestimated it by 31.49 minutes.
Two honest caveats before anyone uses that table as a ranking. The Garmin device tested was a Vivosmart 4, which predates Garmin's current advanced sleep monitoring, so κ = 0.21 is a fact about that band and not about a Fenix 8. And the cohort included people with suspected sleep apnea, where the authors note agreement "tend[s] to decrease as sleep apnea severity increases."
Garmin's own validation is the one people quote least and should quote most. Garmin commissioned a study at the University of Kansas Medical Center, 55 participants sleeping at home against a three-channel EEG reference, and published the result itself: overall algorithm accuracy 69.7%, sensitivity to sleep 95.8%, specificity to wake 73.4%. Garmin telling you its sleep algorithm is right about 70% of the time is the most useful number any of these four companies has published about itself. It was presented as a conference poster rather than in a peer-reviewed journal.
Our Garmin data: the score is real, its predictive value is not
Here is where our own evidence goes, and it is Garmin-only. We have never held Whoop, Oura or Apple data, and nothing in this section should be read as measuring them.
We matched 7,492 steady runs across 54 athletes to the Garmin sleep score from the night before. Median within-athlete correlation between sleep score and running efficiency: 0.04. That is nothing. The full write-up is in our study of whether Garmin's sleep score predicts tomorrow's run, with the quintile breakdown and the per-athlete spread.
The reason I believe the null is the positive control. Run the identical pipeline against Training Readiness, which takes sleep as a direct input by construction, and the same code returns a median correlation of 0.437 with 100% of athletes positive. The method finds relationships that exist. This one does not exist.
The other number worth carrying into a cross-brand argument: Garmin sleep score versus Body Battery came back at a median 0.683 across 85 athletes, again 100% positive. The sleep score and the recovery score are largely one measurement wearing two faces, which is the same structural point we made when we compared Garmin, Whoop and Oura recovery scores.
Cohort caveat, as always: these are self-selected Garmin owners who connected a training-analytics app, so almost certainly fitter and more data-curious than average. Method and exclusions are on our research and methodology pages.
The opinion, stated plainly
I think Apple made the right call and the other three made the wrong one, and the stakes are behavioural.
A sleep score built on stage percentages hands you a number whose largest moving part is the measurement with κ between 0.21 and 0.53. When your deep sleep reads low, you do not know whether you slept badly or whether the algorithm put light sleep in the wrong bucket for 40 minutes. You cannot tell those apart, and neither can the company, because it never told you the weights. So you go looking for a cause that may not exist, sleep worse the following night worrying about it, and the score drops again.
Apple's three inputs are all things a wrist can genuinely measure and a human can genuinely change. Go to bed at a consistent time, stay in bed long enough, reduce awakenings. That is also, as it happens, roughly the whole of evidence-based sleep hygiene.
What to do instead
Use duration and timing, ignore the stage split. Total sleep time and bedtime consistency are the parts every device measures well. Stage percentages are the part every device measures badly.
Track the trend, not the night. A 7-day rolling average of sleep duration survives a single mis-scored night. A single score does not.
Never compare scores across brands. A Garmin 78 and an Oura 78 are different quantities computed from different inputs on different scales. Comparing them is a category error, in the same way that comparing Garmin and Apple Watch training metrics head-to-head mostly compares two philosophies rather than two measurements.
Do not plan training off last night's sleep score. Our 0.04 is Garmin-specific, but the mechanism is not: no wrist device has published evidence that its sleep score predicts next-day performance. If you want a load decision, use load. Our write-up on what sleep tracking is actually good for in athletes covers the cases where it does earn its place.
Frequently Asked Questions
Which sleep tracker is most accurate?
For sleep-stage classification against polysomnography, the Apple Watch Series 8 scored highest of the six devices in Schyvens et al. (2025) at κ = 0.53, with Whoop 4.0 at 0.37 and a Garmin Vivosmart 4 at 0.21. For detecting sleep versus wake at all, every device tested in both studies was good at spotting sleep and poor at spotting wake.
Does Garmin publish how the sleep score is calculated?
Partially. Garmin names the three components (duration, quality, HRV-derived overnight recovery) and the signals behind the stages, but publishes no weights, so the score cannot be reconstructed from your own data.
Is Apple's sleep score too simple?
It is simpler on purpose. It contains no sleep-stage or HRV term even though Apple Watch estimates stages and publishes a validation white paper on them. Simple and auditable beats sophisticated and opaque when you are going to act on the number.
Can I compare my Oura score to my partner's Garmin score?
No. Different inputs, different unpublished weights, different scales. The only defensible comparison is your own score against your own baseline on the same device.
Does a bad sleep score mean I should skip my workout?
Not on its own. In our Garmin data, the median athlete's correlation between last night's sleep score and next-day running efficiency was 0.04 across 7,492 runs. Decide on how you feel and on your training load.
There is a version of this comparison that never gets written, because no company will fund it: the same person, four devices, one night, four scores, published side by side with the PSG that settles it. Until someone does that at scale, the most honest thing any of these numbers can tell you is whether tonight looked like your own last month. Every device answers that question adequately. None of them has earned the right to answer a bigger one.