Every flagship smartwatch review published in the last two years has carried the same implied promise: this generation, finally, the sensors are good enough. More photodiodes. A skin-temperature channel. An on-device model that writes your night up in plain English before you've finished the first cup. I believed enough of it to run the obvious experiment — wear the three current flagships simultaneously for twenty-four nights, dose half of those nights with 150 mg of caffeine at four in the afternoon, and see which watch noticed.
All three noticed. None of them noticed in the place they point you to look.
The verdict: the Samsung Galaxy Watch Ultra 2, the Apple Watch Ultra 3 and the Garmin Fenix 8 all registered my caffeinated nights as worse, but the evidence sat in overnight heart rate and heart-rate variability — things the sensor measures more or less directly — and was close to invisible in the sleep-stage breakdown, which is the screen all three of them open to.
The myth worth killing
The myth is no longer "sleep trackers are accurate." Nobody sophisticated says that out loud anymore. The current version is subtler and much stickier: that inaccuracy was a sensor problem, that sensors have improved a great deal, and that the gap is therefore closing on its own.
It's a reasonable inference from the outside. The hardware genuinely has improved, and the trend across premium wearables since 2023 has been relentlessly additive — ECG, bioelectrical impedance, blood oxygen, skin temperature, dual-band GPS, and now a summarising language layer on top of all of it. If the picture is sharpening everywhere else on the device, why would sleep staging be the one channel standing still?
Because almost none of what's been added feeds the staging classifier. But that's mechanism, and mechanism comes third. First, nights.
What I actually did
Twenty-four nights across five weeks. Each morning I flipped a coin: heads meant 150 mg of caffeine at 16:00, delivered as two shots of the same commercial capsule so the dose stayed fixed; tails meant nothing after 10 a.m. Twelve nights each way. Same bed, same room, blackout blinds, no alcohol for the duration.
Apple on the right wrist, Samsung on the left, Garmin on the left forearm about four centimetres proximal to the Samsung — the 51 mm Fenix will not share a wrist with anything. I rotated the arrangement every third night, because the Ultra 3 read consistently differently on my non-dominant side and I wanted that noise spread across both conditions rather than baked into one.
What I don't have is ground truth. I have no polysomnograph, so every staging number below is one algorithm's opinion checked against two other algorithms' opinions. What I do have is a bedside timer I punched at lights-out, a morning self-report of how long I thought I'd lain there, and a Polar H10 chest strap worn on nine of the twenty-four nights as a cardiac reference. Self-reported sleep onset is a poor instrument — people who are anxious about sleep overestimate it, and by week three I was a person who was thinking about sleep constantly. Read it as directional.
What the watches said
| Source | Median total sleep | Median "deep" | Median onset | Caffeine night read as worse |
|---|---|---|---|---|
| My log | — | — | 21 min / 38 min | 11 of 12 |
| Galaxy Watch Ultra 2 | 6 h 52 | 62 min | 14 min | 7 of 12 |
| Apple Watch Ultra 3 | 7 h 11 | 51 min | 9 min | 8 of 12 |
| Garmin Fenix 8 | 6 h 39 | 44 min | 17 min | 9 of 12 |
My log shows control nights first, caffeine nights second. "Worse" means the device's own headline judgment — Samsung's and Garmin's sleep scores, and for the Apple, which offers no single number, lower total sleep plus more time awake.
Start with the disagreement, because it sets the ceiling on everything else. On identical nights, worn simultaneously on the same body, the three watches differed on total sleep time by a median of 31 minutes, with a worst case of 1 h 04. On deep sleep the median spread was 24 minutes — Samsung routinely credited me with roughly 40% more slow-wave sleep than Garmin for the same eight hours. They cannot all be right, and no amount of firmware makes them all right at once.
Then the failure that matters for this particular question. Caffeine's most reliable effect on me is at the front of the night: my own log put onset at a median of 21 minutes on control nights and 38 on dosed ones, a 17-minute gap. The watches saw that gap as three to six minutes. All three also underestimated my absolute onset by seven to twelve minutes, which is the expected direction — lying still with a slow pulse looks like sleep to an accelerometer and a photodiode. Whatever caffeine did to my first half hour, the watches largely slept through it.
The one thing all three got right
Overnight heart-rate minimum. On control nights my floor sat at a median of 48 bpm; on caffeine nights it was 52 to 53 bpm on every device, and the H10 agreed within 2 bpm on the nine nights it was on. Overnight HRV fell 12–18% depending on which watch you asked. Most striking: the nadir arrived a median of 71 minutes later in the night on dosed nights, consistently, across all three brands.
That is a clean, replicable, cross-vendor signal produced by three independent algorithms. It is the most useful thing I got out of five weeks. Not one of the three puts it on the first screen.
Why the stage chart is the wrong instrument
Sleep stages are defined by cortical electrical activity. A wrist device sees none of it. It sees green light scattered off capillaries, from which it derives interbeat intervals; it sees motion; it sees skin temperature. A model trained against lab recordings maps those proxies onto stage labels. It's inference from the autonomic nervous system, not observation of the brain.
Now consider what caffeine is. It's an adenosine receptor antagonist, and adenosine is both the molecule building your sleep pressure and part of what quiets your cardiovascular system overnight. Block it and heart rate stays up, HRV stays suppressed. Which means caffeine directly perturbs the exact inputs the classifier uses to decide light versus deep. To a model trained mostly on ordinary nights, an elevated pulse with flattened variability looks like shallower sleep — whether or not the cortex agrees.
So when a watch tells you the afternoon espresso cost you 19 minutes of deep sleep, some unknowable fraction of that number is caffeine acting on your heart rather than on your slow waves. The number isn't fabricated. It's just measuring a different thing than the label claims, and the label is the part people screenshot.
Meanwhile the additions of the last two generations — ECG, impedance, SpO2, the narrative coaching — mostly sit outside the staging pipeline entirely. The upgrade in premium wearables recently has been the language model, not the photodiode.
Where each one actually wins
Garmin Fenix 8. Eleven days between charges in my test, with nightly tracking and one recorded activity a day. For sleep work that is the specification that matters, because a tracker you charge overnight is a tracker with holes in its dataset. Garmin also surfaces autonomic data most honestly — Body Battery is the closest any of the three comes to leading with the number that actually moved. Against it: the menu system is a filing cabinet, and a 51 mm case leaves a bezel imprint on your forearm. Twice I woke up and took it off.
Apple Watch Ultra 3. The best of the three at the simple sleep-versus-wake question, and the least intrusive — it doesn't nag, and wrist detection just works. It's also the weakest at longitudinal thinking: no single score, trends buried three taps into Health. Two nights per charge means you're topping up during dinner or losing data. For a device this good at collecting, that's the constraint I resented most.
Galaxy Watch Ultra 2. The one I'd have kept if forced to choose one, purely because it's the least present on the wrist at 3 a.m. Its sleep-apnea detection is a genuinely serious feature and I don't want to undersell it. Its AI summaries are well written and confidently quote stage minutes I have no way to verify, which is a strange combination to hold in your hand. And it is meaningfully less useful paired with anything other than a Galaxy phone; that's a real cost, not a footnote.
Who this is for, and who it isn't
If you want to A/B test caffeine on yourself, any of the three works — but read the overnight heart-rate curve and ignore the stage chart. Honestly, a £100 band with the same optical stack gives you that curve just as well; this is the rare case where the premium tier buys you nothing for the job.
If you run trails, Garmin remains the answer, and Samsung still can't build a route in its phone app. If you dive seriously, none of these is your watch — that's a Descent. If you're already inside Samsung's ecosystem and want something you'll wear to bed without noticing, the Ultra 2 is the easy pick.
It isn't for anyone who wants a stage-accurate number, because that product doesn't exist at this price. And if you suspect an actual sleep disorder, a smartwatch review is the wrong reading material — get a real study.
The winner for this specific job is the Fenix 8, on a spec nobody markets as a sleep feature: eleven days between charges is eleven consecutive nights with no gaps.
The myth: the new sensor generation has finally made consumer sleep staging accurate enough to test what your afternoon coffee is doing to you. The more accurate version: the new sensor generation is very good at showing you that your afternoon coffee kept your heart working past midnight — and the stage chart it wraps that finding in is the least trustworthy thing in the report.