Your watch told you that you slept 7 hours and 14 minutes, that your "sleep score" was 87, and that you spent 14% of the night in deep sleep. The number felt good. An 87 sounds like a B-plus, like you did something right.
Here is the uncomfortable part. That 87 was generated by a proprietary algorithm that has, in most cases, never been published, never been independently validated against the clinical gold standard, and never been the same formula from one firmware update to the next. Wearable sleep trackers are now on tens of millions of wrists, and they are genuinely useful. But the score is not a diagnosis, and the question almost nobody asks out loud is the one worth answering: at what point does the data on your wrist stop being a hobby and become a reason to call a doctor?
That is the whole piece. The honest answer involves a lot of "it depends," and I'll show you where.
The question, stated plainly
You bought the thing, or it came with the phone, and now you have months of charts. Resting heart rate trending down — good, probably. A jagged hypnogram with five "awakenings" you don't remember. A blood-oxygen graph that dipped to 91% at 3 a.m. and you have no idea whether that's normal.
The instinct is to either ignore all of it or panic about all of it. Both are wrong, and they're wrong for the same reason: a wearable measures a few things directly, infers a lot of things indirectly, and is much better at trends than at any single reading. To know when to escalate, you have to know which is which.
What your tracker actually measures, in the order it happens
Start with what the device can physically detect, because everything downstream is built on it.
First, motion. Every consumer sleep tracker contains an accelerometer. When you stop moving, the device's oldest and most reliable trick — actigraphy — infers that you're probably asleep. Actigraphy has been used in sleep research since the 1990s, and it's decent at one thing: estimating when you were asleep versus awake, as long as you were lying still.
Second, heart rate, read optically. A green LED shines into your skin and a photodiode measures how much light bounces back; blood absorbs more light when a pulse arrives, so the fluctuation gives beat-to-beat timing. This is photoplethysmography, or PPG. From the spacing between beats the device derives heart rate variability (HRV), which tends to rise as you shift into deeper, more parasympathetic-dominated sleep.
Third, on newer devices, blood oxygen saturation (SpO2), estimated by shining red and infrared light and comparing absorption. And skin temperature, respiratory rate (read from the subtle modulation of your pulse as you breathe), and sometimes movement-derived snoring detection through the phone's microphone.
Now the inference. The device does not see your brain. A clinical sleep study — polysomnography — stages sleep using an EEG reading the actual electrical activity of your cortex, plus eye movement and chin muscle tone. Your watch has none of that. So it takes the raw streams it does have — stillness, falling heart rate, rising HRV, slowing breath — and feeds them into a model that outputs a guess: light, deep, REM, awake. The hypnogram you see in the morning, with its tidy colored bands, is a statistical reconstruction. A plausible one, often. But a reconstruction.
That distinction is the entire basis for knowing what to trust.
Are wearable sleep trackers accurate?
It depends entirely on what you're asking them to be accurate about. They are good at detecting whether you're asleep or awake and at tracking how long you slept; they are mediocre at correctly labeling sleep stages; and their single-night blood-oxygen readings are too noisy to diagnose anything on their own.
That's the snippet version. Here's the substance behind it.
The recurring benchmark in this field is sensitivity (how well a device catches the minutes you were actually asleep) versus specificity (how well it catches the minutes you were actually awake). Consumer devices have long scored high on sensitivity and poorly on specificity — meaning they tend to assume you're asleep when you're lying still but awake, and therefore overestimate total sleep. A frequently cited evaluation here is de Zambotti et al. (2019), published in Chronobiology International, which tested a then-current Fitbit against polysomnography and found total-sleep-time estimates within a reasonable margin but a familiar weakness at scoring wake after sleep onset.
Stage-by-stage, the picture is rougher. A multi-device study, Chinoy et al. (2021) in Sleep, compared seven consumer trackers against in-lab PSG and against research-grade actigraphy. The headline: most did fine on total sleep time but disagreed meaningfully with the gold standard on how that time was divided into light, deep, and REM. Deep-sleep estimates, in particular, are where I'd put the least faith. If your watch says you got "only 9% deep sleep," that may say as much about the algorithm's assumptions as about your brain.
So when someone asks whether these devices are accurate, the honest reframing is: accurate enough to be a useful longitudinal signal, not accurate enough to be a verdict. Two things follow. A single night's data is close to meaningless. A two-month trend is where the value lives.1
A note on honesty about my own position: I'd tell you to do your own research on your specific device, because "wearable sleep trackers" is not one thing. A 2024 flagship watch with an FDA-cleared sleep-apnea-notification feature and a $30 fitness band running a five-year-old algorithm are not in the same conversation, even though both will happily print you a sleep score.
The difference between a bad night and a signal
Everyone has bad nights. The art of using this data well is learning to ignore them.
Here is the threshold I'd suggest, and I'll defend it: one off reading is noise; a pattern that persists for two to four weeks and lines up with how you actually feel is a signal. The "lines up with how you feel" part matters because the device can be wrong, but your daytime function — the unshakable 2 p.m. fog, the morning headaches, falling asleep in meetings — is a real symptom that a clinician will take seriously regardless of what your watch says.
A few patterns are worth more attention than others.
Recurring oxygen desaturation. A single 91% dip is within the range of normal nocturnal variation, and PPG-based SpO2 on the wrist is notoriously sensitive to fit, skin tone, and motion. But repeated, patterned dips — a sawtooth graph that drops and recovers over and over through the night — is the kind of thing that can shadow obstructive sleep apnea, where the airway collapses, oxygen falls, and a micro-arousal yanks you back to breathe. The wearable can't diagnose apnea. The pattern can be the reason you finally get tested. Some newer devices now offer a "breathing disturbances" or sleep-apnea notification feature precisely because the longitudinal SpO2 trace is suggestive even when any single number isn't.
Heart rate that won't come down. During healthy sleep, your heart rate should descend below your daytime resting baseline, often reaching its low point in the early-morning hours. A nocturnal heart rate that stays elevated, night after night, alongside suppressed HRV, can track with poor sleep quality, alcohol, illness, stress, or overtraining. None of those is a diagnosis. All of them are worth noticing if the elevation is new and durable.
Sleep timing chaos. This one the device measures well, because it's mostly actigraphy — onset and offset, the thing wearables are best at. If your tracker shows your sleep midpoint sliding two or three hours later across the week and snapping back on Monday, that's social jet lag, and it has real metabolic and mood consequences. You don't need a doctor for that. You need a more boring schedule, which is harder.
Notice the spread there. The same dashboard contains one metric the device measures reliably (timing), one it measures decently as a trend (heart rate), and one it can only hint at (oxygen). Treating them all with the same confidence is the most common mistake people make with this data.
What the data is actually good for in a doctor's office
Say the trend is real. You've got three weeks of sawtooth SpO2, you wake unrefreshed, your partner says you snore and sometimes stop. What does the watch do for you now?
Less than you'd hope, and more than you'd think.
What it does not do: diagnose. No clinician will diagnose sleep apnea from a consumer wearable's graph. The diagnostic pathway is a home sleep apnea test or an in-lab polysomnogram, which measures airflow, respiratory effort, and oxygen with validated medical sensors and reports an apnea-hypopnea index — the number of breathing interruptions per hour. Five to fifteen is mild, fifteen to thirty moderate, above thirty severe. Your watch does not produce a real AHI, whatever a "score" implies.
What it does do is change the conversation. Primary-care visits are short, and "I think I sleep badly" is easy to wave off. "Here are three weeks of recurring overnight oxygen dips, my resting heart rate is up 8 beats, and I'm exhausted by mid-afternoon" is a specific, time-stamped complaint that gives a clinician something to act on — most usefully, a reason to order the actual test. The wearable's job is to get you through the door of the right diagnostic, not to replace it.
Two cautions before you go.
First, bring patterns, not screenshots of your worst night. Show the trend line, ideally over weeks. A clinician can do something with "this is my typical month." A single dramatic hypnogram invites them to dismiss the device, fairly.
Second, the condition the literature has actually named: orthosomnia. Coined in a 2017 case series by Baron et al. in the Journal of Clinical Sleep Medicine, it describes patients whose pursuit of "perfect" tracker sleep made their sleep worse — people lying in bed anxious about their score, which is a recipe for the exact insomnia they were trying to fix. The lesson isn't to throw out the device. It's that the data is a flashlight, not a judge. If checking your sleep score is itself keeping you up, the most evidence-based move is to stop checking it for a while.
An honest rule of thumb you can use tonight
Stop reading your sleep score as a grade. Read your data as a question. A single night tells you nothing; ask instead whether a pattern has held for at least two weeks and matches how you feel during the day. If both are true, take the trend — not one screenshot — to a clinician and ask them to order the real test. If neither is true, the most useful thing your tracker can do tonight is stay on the nightstand while you go to sleep without looking at it.
And here's the quick triage, by metric and by how much you should trust it:
| What you see | How much the device knows | What to do |
|---|---|---|
| Total sleep time, bedtime drift | Good — mostly actigraphy | Trust the trend; fix your schedule before you call anyone |
| Sleep stages (deep/REM %) | Weak — inferred, varies by algorithm | Don't optimize for it; watch only big, sustained shifts |
| Resting/nocturnal heart rate, HRV | Decent as a trend | Note new, durable changes; context (alcohol, illness) matters |
| Repeated overnight SpO2 dips + daytime fatigue + snoring | Suggestive, not diagnostic | This is the escalate signal — bring weeks of data to a doctor |
| One alarming reading, otherwise fine | Probably noise | Wait a week before reacting |
The thread running through all of it: the more a number depends on the device inferring something it can't directly see, the less weight you should give any single reading, and the more you should rely on the shape of the trend.
Back to that 87
Return to the score we started with. An 87 out of 100, 14% deep sleep, seven hours and change. We can now say what that number is and isn't.
It isn't a grade, because there's no validated, published rubric behind it, and the deep-sleep figure feeding it is the metric the comparison studies trust least. It isn't a diagnosis, because the device never saw your brain or your airway. And it isn't comparable to last year's 87, because the algorithm probably changed underneath you.
But it isn't nothing, either. If that 87 is part of a stable line stretching back two months, it's a reasonable proxy for "your sleep is roughly consistent" — which is genuinely worth knowing. And if that line breaks — if the scores slide while your afternoons go gray and your overnight oxygen graph turns to teeth — then the most valuable thing the number ever did was give you a date-stamped reason to make an appointment. The score was never the answer. It was the prompt to go find one.
-
A practical consequence: the first week with any new tracker is the least informative week you'll ever have with it, both because there's no baseline yet and because the novelty of being measured tends to change how people sleep. Give it a month before you read into anything. ↩