For ten weeks this spring I slept wearing four sleep trackers at once, because I wanted an answer to a narrow question: which of them, if any, is actually watching my circadian rhythm, and which are just grading the night after it's over. An Oura ring on my left index finger. A Whoop band on my right upper arm. An Apple Watch on my left wrist, running its own sleep tracking. A Withings mat under the mattress on my side of the bed. Sixty-eight usable nights out of seventy — I lost two to a ring I forgot to charge.
The verdict, stated plainly so you can stop reading here if you want: none of the four measures your body clock, three of them estimate its timing by averaging your own sleep history back at you, and the single most useful number I found in ten weeks is one that no device puts on its home screen — the shape of the temperature curve in the first hour after lights out.
I ran this because of a specific irritation. My scores were good. Oura averaged 84, Whoop's sleep performance averaged 88, and I was waking up most mornings feeling like I'd been assembled from spare parts. That mismatch is common enough that it's become the standard complaint about sleep tech, and the standard answer — your sleep timing is off — is true but useless, because nobody tells you which feature on which device would let you see it.
I tested them in the order the body works
Comparing feature lists gets you nowhere; every device claims everything. So I did it differently. Sleep is a chain of events with a fixed order: light enters the eye, the clock takes its instruction, the body starts shedding heat, adenosine pressure tips you over, the night runs its architecture, and you wake into a hormonal ramp. I went station by station and asked, at each one, which of the four devices touches the physiology and which is guessing.
Alongside the devices I kept a paper log, which matters, because a log you keep on your phone becomes a log you keep after checking your score. Every night: lights-out time, my own guess at how long I lay awake, milligrams of caffeine and the clock time of each dose. Every morning, before opening a single app: how I felt, one to ten. My baseline caffeine was about 130 mg of home espresso before eight, plus a 160 mg drip coffee in the afternoon on days when the afternoon deserved it.
Station one: light gets in
This is where the whole chain starts and where the hardware is thinnest.
Of the four devices, exactly one senses light: the Apple Watch, whose ambient sensor feeds the Time in Daylight metric. The ring can't — it lives on a finger that spends the day inside a sleeve, a pocket, or a keyboard's shadow. The band can't. The mat, obviously, can't. So the input that sets your clock is the input your tracker is least equipped to see.
I bought a cheap lux meter to find out what I was actually getting, and the numbers embarrassed me. My kitchen at 6:50 a.m. with every light on: 210 lux. Ten minutes on the pavement outside on a flatly overcast morning: 9,400 lux. A bright one: over 60,000. I had been treating "I get up and it's bright in here" as a morning light dose. It is roughly two percent of a dull walk to the corner.
Evenings ran the other way. My living room after nine reads about 45 lux, which is fine, but the kitchen overhead I stand under while washing up reads 320 lux — brighter than my morning kitchen, at the exact hour I want the opposite. No tracker in the pile flagged that. Time in Daylight came closest to being useful and it counts minutes above a threshold, not lux, and counts them at a wrist that is frequently in a jacket. It told me I got 41 minutes of daylight on a day I spent almost entirely indoors, because I'd had lunch by a window.
What I'd want and nobody sells me: a light log with the same resolution my heart rate gets. Until then, the honest move is a five-dollar phone app and a week of curiosity about your own rooms.
Station two: the clock takes its instruction
The lab measure of circadian phase is dim-light melatonin onset — saliva samples, hourly, in a dim room, under supervision. No ring does that. No band does that. Nothing you can buy does that, and any feature naming itself after your body clock is working from proxies.
Which proxies? Mostly your own behaviour: the rolling average of your sleep midpoint, your wake-time consistency, and — this is the more interesting one — the clock time of your overnight heart-rate minimum. That nadir tends to sit in a reasonably fixed relationship to the rest of the circadian machinery, and it's one of the few phase-adjacent things a wrist or finger can genuinely see rather than infer.
Oura's body-clock feature, on the firmware I had in April, gave me a chronotype and a suggested sleep window. Over 68 nights that window moved by a total of twenty minutes. I genuinely cannot tell you whether that reflects an admirably stable physiology or a slow-moving average of my own habits being handed back to me with confidence, and the app doesn't offer me any way to distinguish those two possibilities. That's not a knock on the estimate; it's a knock on the presentation. A number that can't be wrong in any way I can detect isn't telling me much.
The nadir timing did move, though, and it moved for a reason — which is the next station but one.
Station three: the body has to cool
To fall asleep, your core temperature has to drop, and the way it drops is by dumping heat through your hands and feet. Distal skin warms as core cools. A ring sits on a finger, which makes it, almost by accident, the best-placed consumer device for the most mechanically important thing that happens at bedtime.
This is where the ten weeks paid off. I stopped reading the nightly temperature deviation as a number — that number is for illness detection — and started reading the slope in the first hour after lights out.
On nights where my distal temperature rose by at least 0.3 °C within 45 minutes of lights out, my next-morning subjective score averaged 7.4. On nights where the rise lagged past 90 minutes, it averaged 5.6. Sleep scores on those two groups of nights differed by less than three points. The device knew something the score threw away.
Caveats, and they're real ones. Sixty-eight nights is one person and one season. I chose my own bedtimes, so on nights I went to bed genuinely sleepy I both cooled faster and felt better, and I can't separate those. And I read the slope off a chart with my eyes rather than exporting the data, which means my thresholds are tidier than my method.
What I changed on the strength of it: a hot shower ninety minutes before lights out rather than immediately before, and the bedroom at about 18 °C instead of the 20.5 °C I'd drifted into over winter. The slope got steeper. So did the mornings.
Station four: adenosine, and the argument I lost with my own coffee
JavaSleep, so: caffeine. I ran two fourteen-night blocks. Block A had a hard 10:30 a.m. cutoff. Block B added a 160 mg drip coffee at 2:30 p.m. and changed nothing else — same bedtime target, same room, same shower timing.
Here is what the trackers said:
- Oura sleep score: 85 (A) vs 83 (B)
- Whoop sleep performance: 88 (A) vs 87 (B)
And here is what was underneath:
- Wake after sleep onset: 19 min (A) vs 41 min (B)
- Time from sleep onset to first deep bout: 22 min (A) vs 51 min (B)
- Overnight heart-rate nadir: about 40 minutes later in B
- My own morning one-to-ten: 7.1 (A) vs 5.9 (B)
A two-point score difference is nothing. It's within the noise of any given Tuesday. But a doubling of time-to-first-deep-bout and a nadir shoved 40 minutes later is a legibly different night, and I felt every bit of it. Caffeine's half-life sits around five hours in most adults, which means that 2:30 p.m. cup is still working at something like 80 mg at 7:30 p.m. and 40 mg at half past midnight. It doesn't stop you sleeping. It taxes the front of the night, which is where the deep sleep lives, and then it hands you a score that says you were fine.
This is the clearest case I found of a real effect surviving all the way through the sensors and then being averaged to death at the last step.
Station five: the night itself, where staging falls apart
I don't have a polysomnograph, so I can't tell you which device staged my sleep correctly. What I can tell you is how much they disagreed with each other about the same night, which is its own kind of evidence.
Across 68 nights, ring and band differed on deep sleep by a median of 31 minutes. On the worst night, the ring said 48 minutes of deep sleep and the band said 103. REM was worse in percentage terms and slightly better in absolute minutes. Total sleep time, by contrast, agreed within about twelve minutes on most nights, and all four devices agreed on the broad strokes of when I was in bed and when I got up.
That's the shape of the technology. Sleep versus wake, from movement and heart rate, is a problem these devices have largely solved. Assigning the sleep to named stages is a modelling exercise, and the models disagree because they're different models, not because my brain did four different things.
The failure mode worth knowing about is stillness. One night I lay in near-dark reading a paper book for about half an hour, barely moving. The ring credited me with 24 minutes of sleep. The mat did not — it watches breathing and body pressure, and a person holding a book breathes like a person holding a book. Over the ten weeks the mat's onset estimates landed closest to my handwritten guesses, off by about seven minutes on average, while the ring consistently thought I'd fallen asleep sooner than I had.
Station six: the morning score
A sleep score is a weighted composite, and duration carries most of the weight in every implementation I've looked at. Efficiency and timing regularity take most of the rest. That's a defensible design — duration is the metric with the strongest evidence behind it and the lowest measurement error — but it means the score is structurally incapable of telling you what you came to it for.
My single worst morning of the experiment scored 91. I'd slept 7 h 52 m, in bed at a consistent time, with high efficiency, after a late coffee and a warm room. Every input the score cares about was excellent. Every input it doesn't collect was bad.
The scores aren't lying. They're answering a different question than the one I'm asking when I open the app at 6 a.m.
The grid
| Device | What it genuinely measures | What it infers (and how well) | The one feature worth the money |
|---|---|---|---|
| Oura ring | Distal skin temperature, HR, HRV, movement | Sleep stages (shaky), body-clock timing from your own history (unfalsifiable), onset (optimistic) | The nightly temperature curve, read as a slope |
| Whoop band | HR, HRV, respiratory rate, movement | Stages (shaky, and disagrees with the ring), recovery (a genuinely useful load metric) | Strain-versus-recovery over weeks, not nights |
| Apple Watch | Ambient light, HR, movement | Stages (roughly), daylight exposure (crudely) | Time in Daylight — the only light data in the pile |
| Withings mat | Body pressure, respiration, movement, snoring | Stages (shaky, like everyone) | Honest sleep-onset timing, because it can't be fooled by stillness |
Who this is for, and who it isn't
Buy a ring if you want the temperature curve and you're willing to ignore the score printed above it. Fingers are the right real estate for the one physiological signal in this whole category that's both well-measured and directly actionable at bedtime.
Buy a band if your actual question is about training load and how hard you can go tomorrow. It's a good daytime instrument that also happens to sleep with you.
Buy a mat if you've ever suspected your tracker of flattering you, or if you sleep next to someone whose movement pollutes your wrist data. It's the only device here I never caught in an obvious lie.
Don't buy anything if you're chasing a score. The score will go up. You will not feel better. I have ten weeks of data on precisely this.
And a real limitation on all of the above: this is one person, one body, one spring, and one palate for coffee. I couldn't test against a lab. I couldn't blind myself to my own log. Anyone telling you they ran a clean self-experiment on their own sleep is telling you about their optimism, not their method.
The line worth screenshotting
Buy the ring for the temperature curve, ignore its score, and put the money you were going to spend on a second wearable into blackout curtains and a cheaper afternoon.
The deeper thing I came away with is that the whole category is currently strongest at the end of the chain and weakest at the start. It can tell you, with reasonable confidence, how long you were unconscious. It can tell you something real about your temperature and your heart. It cannot see the light that set your circadian rhythm eighteen hours earlier, and that's the input doing most of the work.
So: if you take one thing to bed tonight, take this — when your tracker and your body disagree about how you slept, the tracker is the one estimating, so write down how you feel before you open the app, and let that be the number you try to move.