At 3:40 on a Tuesday afternoon I drank twelve ounces of cold brew — call it 200 mg of caffeine, though nobody behind the counter could have given me the real number. I went to bed at 11:15. I slept, according to the ring on my left index finger, seven hours and eleven minutes. Then I woke up, reached for my phone before I reached for my glasses, and learned that I had scored a 61.
I have been living inside fitness trackers and wearables since 2019: a wrist band, then a chest strap, then a ring, and for one stretch I would rather not defend, two devices at once, on the theory that where they disagreed the truth would be somewhere in the middle. A 61 is not a catastrophe. It is a C-minus. And by every standard my body was able to report that morning, the night had been fine. I woke before the alarm. I made coffee without resentment.
I did not feel fine by ten.
This is the part I want to be precise about, because it usually gets gestured at and never examined. Nothing measurable about my physiology changed between 6:50 and 6:51 a.m. What changed was that I read a two-digit integer. By mid-morning I had located a fogginess I would otherwise have walked straight past, and by two in the afternoon I had built a small theory about it involving the cold brew, my age, and a vague sense that I had been getting away with something for years.
So I ran the smallest experiment I could design that would tell me anything. A friend portioned fourteen paper bags of ground coffee — seven caffeinated, seven decaf, same roaster, numbered and shuffled — and I brewed one at 3:30 p.m. every day for two weeks. Each morning, before I opened the app, I wrote two numbers on an index card: how well I thought I'd slept, out of ten, and which kind of coffee I thought I'd had.
I guessed the caffeinated days correctly eight times out of fourteen. That is a coin flip with a slight limp. My handwritten sleep ratings and the device's scores agreed loosely and disagreed loudly on the tails: my two worst-feeling mornings both scored in the seventies. And the number that best predicted how I rated my energy at 3 p.m. was not the caffeine and not my morning index card. It was the score.
This is n=1, unblinded in every way that matters, with no statistics worth the name. It is not evidence about caffeine. It is evidence about me, and the reason I'm telling you is that I suspect a lot of you have run the same accidental study without writing anything down.
What the number actually knows
Consumer sleep devices are genuinely good at one thing and mediocre at another, and the gap between the two is where most of the trouble lives.
Chinoy et al. (2021), in the journal Sleep, ran seven consumer sleep-tracking devices against in-lab polysomnography in 34 healthy adults. The devices were excellent at noticing sleep — sensitivity above 0.9 for most of them. They were poor at noticing wake. Specificity for wake ran roughly from 0.18 to 0.54 depending on the device, which means that when you were actually lying there awake, most of these devices had worse-than-coin-flip odds of registering it. Stage-level agreement — how much REM, how much deep — was substantially weaker than the sleep/wake call, which is the opposite of the ranking most people assume.
That asymmetry has a physical explanation. A wrist or finger device sees movement and a photoplethysmography signal: light bounced off capillaries, converted into heart rate and beat-to-beat variability. Lying very still with your eyes closed and your heart rate low looks, to that sensor stack, almost exactly like sleeping. Marco de Zambotti's group has spent a decade documenting where these estimates hold up and where they drift, and the honest summary is that they are respectable trend instruments and unreliable single-night verdicts.
Then there's the part nobody discloses. A sleep score is not a measurement. It is a proprietary weighted composite of several estimates, and the weights change. Manufacturers push algorithm updates the way they push firmware, quietly, and your March data and your November data may have been produced by meaningfully different models.1 You are not tracking a trend in your sleep. You are tracking a trend in your sleep as filtered through a black box that someone is still editing.
I don't say this to be dismissive. Fitness trackers and wearables surfaced something real for a friend of mine: a hard, undeniable log showing she had gone to bed after 1 a.m. eleven nights out of fourteen while insisting she was "basically an eleven-thirty person." No amount of introspection was going to produce that. The instrument was right and she was wrong, and she moved her bedtime.
Why does a bad sleep score make the day worse?
Because your expectations about how you slept influence how you function almost as much as the sleep itself does — and when researchers hand people fabricated sleep data, the fabricated data moves the outcome.
The cleanest demonstration is Draganich and Erdal (2014), in the Journal of Experimental Psychology: Learning, Memory, and Cognition. Undergraduates were hooked up to equipment, given a short lecture on how REM sleep supports cognition, and then told a specific figure for the percentage of REM they'd gotten the night before. The figure was invented. Participants told they'd had above-average REM did measurably better on the PASAT, an auditory arithmetic task, and on a verbal fluency measure. Those told they'd had below-average REM did worse. Nobody's actual sleep differed. Two experiments, undergraduate samples in the fifties, effect sizes moderate.
Gavriloff et al. (2018), in the Journal of Sleep Research, ran a sharper version in people who actually had insomnia. Participants were given sham feedback from an actigraphy device — told, regardless of what the device recorded, that they'd slept well or slept badly. Negative feedback increased daytime sleepiness ratings, worsened mood, and — the finding I keep returning to — increased self-reported monitoring for sleep-related threat. The sample was small, under twenty. This is a suggestive result, not a settled one.
Which is the honest characterization of this whole literature: repeatedly demonstrated in small rooms with small samples, plausible on mechanism, and not the same thing as established in the wild over months of real use. The data here is thinner than the confidence with which it usually gets cited, including by people making my argument.
The loop, in the order it happens
Here is what a 200 mg cup at 3:40 p.m. and a score of 61 the next morning actually do, in sequence.
3:40 p.m. — absorption
Caffeine is absorbed nearly completely from the small intestine and reaches peak plasma concentration in about 30 to 60 minutes. It crosses the blood-brain barrier easily because it is structurally similar enough to adenosine to sit in adenosine's receptors — principally A1 and A2A — without activating them. Adenosine has been accumulating in your brain since you woke up, a byproduct of the ATP your neurons have been spending. That accumulation is most of what sleep pressure is. Caffeine does not drain the tank. It puts tape over the gauge.
By 8 p.m. — clearance, mostly
Caffeine is metabolized in the liver, overwhelmingly by the enzyme CYP1A2. The population median half-life is about five hours, but the range is embarrassing for anyone who wants to give general advice: roughly 1.5 to 9.5 hours in healthy adults. Smoking induces CYP1A2 and can nearly halve it. Oral contraceptives inhibit it and can roughly double it. Pregnancy pushes it far longer still. So of that 200 mg, somewhere around 100 mg is still circulating at bedtime for a median metabolizer, and about 50 mg at three in the morning — a shot of espresso, quietly, in the middle of the night.
11:15 p.m. — the night itself
The canonical citation is Drake et al. (2013), Journal of Clinical Sleep Medicine, which gave 12 subjects 400 mg of caffeine at 0, 3, and 6 hours before bed and measured with polysomnography. The 6-hour dose still reduced objective sleep by more than an hour. Twelve people is twelve people, but the direction of that finding has held up.
The subtler effect is on sleep architecture. Landolt and colleagues showed in the mid-1990s that caffeine suppresses slow-wave activity — the low-frequency EEG power that indexes the depth of non-REM sleep — and that a 200 mg dose taken in the morning was still detectable in the nighttime EEG spectrum. So you can sleep a normal number of hours and sleep them shallower.
Meanwhile, adenosine blockade raises sympathetic tone. Caffeine acutely increases blood pressure by a few mmHg and tends to reduce heart rate variability. Hold that thought.
6:50 a.m. — the score
Your device did not observe your slow-wave activity. It observed movement and beat-to-beat cardiac timing, and it inferred everything else. Heart rate variability is one of the heaviest inputs to most sleep scores. Caffeine reduces heart rate variability. It is therefore entirely plausible that a portion of your bad score is the device reading caffeine's cardiac signature rather than the quality of your sleep — a shadow of the drug, not of the night. Plausible, mechanistically tidy, and as far as I can find, not directly tested. I'd like someone to test it.
6:51 a.m. — the reading
And now the number enters a nervous system that has been awake for ninety seconds and is looking for something to do. Harvey's (2002) cognitive model of insomnia, in Behaviour Research and Therapy, describes the mechanism plainly: a person worried about sleep begins selectively monitoring for evidence of impairment, finds it — the body always supplies ambiguous evidence — interprets it as confirmation, and generates the physiological arousal that then interferes with sleep. The model was written before consumer sleep tracking existed. It reads like it was written about consumer sleep tracking.
Health anxiety has a shape, and the device fits it exactly
The cognitive-behavioral account of health anxiety, worked out largely by Paul Salkovskis and colleagues from the late 1980s onward, turns on a specific and well-replicated mechanism: checking and reassurance-seeking produce short-term relief and long-term escalation. The relief is real. It also decays, and the decay teaches you that the checking is what made the anxiety stop, which raises the frequency of the checking. Salkovskis called these safety behaviors, and their defining property is that they prevent you from ever discovering that the feared outcome wasn't going to happen anyway.
A sleep score is reassurance with an expiration date of one day. That is an unusual structure. A blood test reassures you for months. A device that regenerates a fresh verdict every morning cannot reassure you for longer than a morning, and it guarantees that some mornings the verdict will be bad, because scores are constructed to have variance. If they didn't vary, you'd stop opening the app.
Baron et al. (2017), in the Journal of Clinical Sleep Medicine, gave this a name: orthosomnia — patients arriving at a sleep clinic whose presenting complaint was their tracker data, some of whom spent excessive time in bed trying to improve their numbers, several of whom held to the device's interpretation over the clinician's. It is worth saying clearly what that paper is: a small case series, no control group, no prevalence estimate. It is the weakest tier of clinical evidence and by a wide margin the most-quoted item in this entire conversation, including in this article. Both things are true.
The adjacent finding, which the common citation attributes to work on wearable use in patients with atrial fibrillation, is that patients who used consumer devices to monitor their condition reported more symptom-related anxiety and made more unscheduled contact with their clinicians than non-users. That research is cross-sectional. It cannot tell you whether monitoring produced the anxiety or whether anxious patients bought more monitors. This is the central unresolved problem in the field, and no amount of confident writing about wellness culture makes it go away.
What the psychology does predict, fairly specifically, is who is at risk. Not everyone. The vulnerability markers are high intolerance of uncertainty, a pre-existing tendency toward bodily checking, and — most relevant here — a history of using numbers as a way to manage feelings. If none of those describe you, a sleep score is probably just an interesting graph.
What's solid, what isn't
| Claim | Status |
|---|---|
| Caffeine within ~6 hours of bed measurably reduces objective sleep | Well established (Drake 2013, PSG-confirmed) |
| Devices detect sleep well and wake poorly; staging is weakest | Well established (Chinoy 2021, 34 adults, 7 devices) |
| Being told you slept badly worsens next-day mood and performance | Plausible; replicated only in small lab samples |
| Caffeine's HRV suppression inflates "bad night" scores | Mechanistically plausible; not directly tested |
| Sleep tracking causes clinical insomnia in vulnerable people | Case reports only; no prevalence data, no controls |
| A sleep score should determine your caffeine plan for the day | No supporting evidence whatsoever |
An honest rule of thumb
Tonight, before you open the app tomorrow morning, do this: write a number from one to ten on a piece of paper for how you slept. Paper, not phone — the phone is where the score lives. Then open the app.
Do it for two weeks and look at the two columns. If your ratings and the scores mostly agree, the device is telling you something you already knew, which is fine and mildly useless. If they disagree, you now have to decide which one you're going to let run your day, and you should make that decision in the abstract rather than at 6:51 a.m. with a 61 on the screen.
Two smaller rules I've kept. Don't look at the score before noon; by then the day has supplied its own evidence and the number can only confirm or contradict it, not create it. And once a week, take the thing off overnight. Notice what that costs you. If taking it off produces a small spike of unease — if you find yourself thinking the night won't count — the device has stopped being an instrument and become a ritual, and rituals are not improved by better sensors.
The thing I still can't answer
I'm still wearing the ring. I want to be honest about that, because the tidy version of this essay ends with me dropping it in a drawer, and the tidy version would be a lie. What changed is smaller. I stopped treating the score as an input and started treating it as a record — something to read at the end of the month, not the start of the day. I moved my last coffee to two o'clock, which is the one intervention here with actual polysomnography behind it. My scores did not noticeably improve. My mornings did.
But the question underneath all of this is genuinely open, and the confident essays on either side are both overreaching. Nobody has run the trial that would settle it: take several hundred adults, measure health anxiety and sleep with objective instruments at baseline, randomize half to a consumer tracker and half to nothing for twelve months, and see which group ends up more anxious and which ends up sleeping better. It would be expensive and awkward — you cannot blind someone to a ring on their finger — and until someone does it, every argument about orthosomnia is an argument about direction of causation supported by case reports and cross-sectional surveys.
So I don't know. After fourteen nights of blinded coffee and six years of graphs, I can't tell you whether the anxious sleep worse because they measure, or measure because they sleep worse. Which leaves the only question I actually care about: is the ring measuring my sleep, or has it become one of the things it measures?
-
This is not a conspiracy claim; it's a product-development fact. Several manufacturers have publicly revised their staging algorithms after validation studies criticized them. The problem is that longitudinal self-comparison — the entire premise of a personal sleep trend — quietly assumes a stable instrument, and you have no way to know when yours changed. ↩