I spent thirty nights trying to prove that the most expensive part of my wearable device ecosystem was also the least useful part. It is. The sensors are quietly good. The subscriptions layered on top of them are streaming services in nicer packaging: you aren't buying a measurement, you're renting an interpretation, and the interpretation can be recut, repriced, or switched off while you sleep through it.
The verdict: every free tier I tested caught the thing I was testing for — a four-hour shift in my caffeine cutoff — and every paid tier mostly restated that finding in friendlier typography. Exactly one paid feature earned its money, and it was not a sleep score.
That is a large claim for one person, three gadgets, and a coffee habit. The rest of this is me earning it.
The setup
From 12 July to 10 August I wore three devices simultaneously, every night: an Oura ring on my left index finger, an Apple Watch on my left wrist, a Garmin Venu on my right. Charging happened in a fixed forty-minute window after dinner so nothing missed a night. Thirty nights, no gaps, no travel, one bed.
Bedtime was 22:45 give or take twenty minutes. Bedroom held at 18°C. No alcohol for the full month, which is less noble than it sounds — alcohol mangles sleep architecture badly enough to swamp a caffeine signal, and I only had one variable to spend.
The coffee was constant: one 5 lb bag of a washed Colombian, 22 g dosed into 360 g of water, pour-over, which by the roaster's own lab figure lands near 200 mg of caffeine per cup. Two cups a day, every day, all thirty nights. Same beans, same grind setting, same water.
None of the three platforms will speak to the others in any useful way, so the comparison happened in a spreadsheet I filled in by hand each morning — about six minutes a day, call it three hours over the month. That is the tax the ecosystem charges you for owning more than one brand, and no subscription I could find removes it.
The only thing I changed was when the second cup happened.
One variable, thirty nights
Alternating blocks of five: five nights with the last cup at 13:30, five nights with it at 17:30, repeated three times. Fifteen nights per condition. Blocks rather than coin flips because I have a job, and 13:30 coffee is a decision you make at breakfast.
The design has an obvious hole and I won't pretend otherwise. I knew which block I was in. No blinding, no placebo decaf, n=1. If I expected worse sleep after a 17:30 cup, I had every chance to deliver it.
Averaged across each condition, the devices said:
- Sleep onset latency: 11 minutes early-cutoff (range 6–19), 24 minutes late-cutoff (range 9–48)
- Wake episodes: 2.1 early, 3.6 late
- "Deep" sleep: 71 minutes early, 58 minutes late
A thirteen-minute latency penalty for moving one cup four hours later. Nothing exotic.
I also logged the clock time when I put the book down, plus my own guess at how long I then lay there. My guesses tracked the hardware at roughly the accuracy of a coin toss weighted toward optimism. On the worst late-cutoff night I would have told you twenty minutes; the three devices said 41, 44, and 48. Whatever else is wrong with these things, they are better at this than I am.
Three platforms, one shift
| Platform | What I paid (Aug 2026) | Caught the caffeine shift? | What I could take with me |
|---|---|---|---|
| Oura ring | $5.99/mo membership | Yes — latency +14 min | Daily summaries as CSV; per-epoch detail flattened |
| Apple Watch | $0 | Yes — latency +11 min | Full Health XML, per-night stage records |
| Garmin Venu | $0 (declined Connect+ at $6.99/mo) | Yes — latency +13 min | FIT files plus a full account archive |
Apple does not hand you a latency figure. I derived it from the gap between the in-bed timestamp and the first asleep record in the exported XML, which took an afternoon of squinting at a schema and a short script. Garmin's came out of the FIT files the same way. Both were free and both were work; the paid tier's real product is that somebody has already done the squinting for you.
The stage numbers are where the honesty gets uncomfortable. On night 14 the ring reported 82 minutes of deep sleep. The watch reported 41 for the same night, same body, same eight hours. Neither hedged. Both drew a confident colored bar underneath.
So stop reading these as measurement devices. They are delta machines. All three disagreed wildly on absolute minutes and all three agreed on direction and near-agreed on magnitude for latency — 11, 13, 14 minutes. Three vendors, three sensor sites, one consistent answer to the only question I actually asked.
I could not test Whoop. It has no free tier to compare a paid tier against, which is, in its way, the finding.
Where the money actually went
Here is the paid feature I would buy again: the chart.
Oura without a membership shows me three numbers. Oura with a membership shows me the plot underneath them — the per-night hypnogram, the thirty-day trend, the heart rate curve across the whole night. The numbers are an opinion. The plot is closer to evidence, and the plot is what showed me that my late-cutoff nights weren't uniformly worse. They were unremarkable after about 02:00 and ugly before it. A single score of 74 cannot tell you that. A chart tells you in one second.
What I would not pay for again at any price is the score. It is a weighted sum of numbers I can already see, with the weights undisclosed, and it compresses out the only part worth looking at.
I declined Garmin's Connect+, so I can't tell you whether its AI summaries beat the free charts. I can tell you the free charts had already answered the question I brought.
The cancellation test
On night 31 I cancelled the Oura membership to find out what the streaming comparison was worth.
The scores stayed. The charts went. Thirty nights of finger-temperature and heart-rate curves that my own body had generated, still sitting on a server, now behind a paywall I had just stepped outside of. Nothing was deleted. It was delisted, in exactly the way a show you were halfway through leaves a catalogue.
Export flows vary and change often enough that you should check yours rather than trust mine, but the pattern held across all three: the summaries are portable, the resolution is not. You can leave with your daily numbers. You generally cannot leave with the raw shape of the night that produced them.
There is a second version of the same problem. Twice in two years a staging algorithm has been updated underneath me and my history changed shape overnight — nights I remembered as bad were quietly promoted. That is a recut, applied to my own past, and nobody asked.
Who this is for, and who it isn't
If you own the hardware and want to answer one specific question — does the afternoon cup cost me sleep, and how much — the free tier is enough, and it was enough on every platform here. Run blocks, watch the delta, ignore the score.
Buy the subscription if you want the charts and you will genuinely open them. That is a real product and I still look at mine several times a week.
Don't buy it for a number that grades your night; you were present for the night. And don't buy it expecting to buy your way out of the lock-in. More money buys a better view of your data inside one company's walls, not a copy of it in your hands. Export first, then decide what the view is worth.
The verdict
My last cup is at 13:30 now, permanently. I learned that for nothing, on three devices that disagreed about almost everything except the part that mattered.
Pay for the chart. Never pay for the score.