The notification arrived at 11:07 p.m. It said my sleep score had been trending down and suggested I consider winding down earlier. I was, at that moment, winding down. Then a phone lit up in my hand to tell me about it.
The latest Google Health app update does something about this, and it is smaller and more useful than a press release would make it sound: on recent Android and iOS builds, you can collapse or hide the AI Coach card from your home feed. Settings → Home screen → manage cards, on most builds. The toggle is server-flagged, so having the current version does not guarantee you will see it — a fair number of people on the app version with the feature still have the old feed, which is standard for Google's staged rollouts and mildly infuriating anyway.
What the toggle does: it removes the coach's daily narrative from the surface you look at first. What it does not do: stop the underlying data collection, cancel your Fitbit Premium billing, or turn off the model that generates the advice. You are hiding a card, not opting out of a system. That distinction matters, and Google has not been especially loud about it.
Why a sleep magazine cares about a UI toggle
Because the AI Coach's most confident, most frequent, most emotionally sticky output is sleep advice — and sleep is the domain where consumer wearable data is weakest. The coach speaks in the register of a clinician reviewing a chart. The chart is an inference built on a wrist.
That gap is the whole story.
What your watch actually measures while you sleep
Walk it through in the order it happens in your body.
You lie down and stop moving. The accelerometer registers that — this is actigraphy, and it is the oldest and most reliable signal in the stack. Sleep versus wake is mostly a movement question, and wrist devices have been decent at it since the 1990s.
Then a green LED on the underside of the device fires into your skin and a photodiode measures how much light comes back. Blood absorbs green light; each heartbeat changes the absorption. That is photoplethysmography, PPG, and it gives the device a heart rate and, from the variation between beats, a heart-rate variability estimate.
As you fall asleep, parasympathetic tone rises. Heart rate drops, often by 10 to 20 beats per minute from your daytime resting figure. HRV typically climbs. Breathing rate steadies. In REM, heart rate becomes erratic again and motion suppression is near-total.
A classifier reads those three streams — movement, heart rate, beat-to-beat variability — and assigns each 30-second window to light, deep, REM, or awake. It has never seen your brain. Sleep staging in a laboratory is defined by EEG, by electrical activity at the scalp, plus eye movement and chin muscle tone. Your watch is inferring a neurological state from a cardiovascular proxy at the wrist. Sometimes it is right. It is not measuring the thing it names.
Three ways to answer "did I sleep well?"
Grade them on three things: how well they track against polysomnography, the lab standard; how actionable the output is; and what they cost you in attention and anxiety.
The AI Coach narrative. A paragraph of generated prose each morning telling you what your night meant and what to do about it. Accuracy: inherits every limitation of the staging underneath it, then adds a layer of causal language the data does not support. Actionability: high on the surface — it gives you an instruction — but the instructions regress to the mean of sleep-hygiene advice you already know. Attention cost: highest of the three, and not only because of reading time. It arrives with authority, and authority about a number you cannot verify is a specific kind of expensive.
The raw metrics, same app. Total sleep time, wake episodes, resting heart rate, the trend line over weeks. Accuracy against PSG: this is where the honest evidence lives. Chinoy et al. (2021), in SLEEP, ran seven consumer devices against polysomnography in 34 healthy adults across multiple nights. The headline was consistent with a decade of similar work — the devices did reasonably at discriminating sleep from wake, and poorly at four-stage classification, with deep-sleep estimates in particular drifting well off the lab measurement. Total sleep time was usable. Stage breakdowns were not. Actionability: moderate, and it improves as you look at more nights. Attention cost: moderate.
A paper diary and a fixed wake time. Bedtime, wake time, one line about how it felt, one line about your last caffeine. Accuracy: subjective sleep-onset estimates are famously imprecise — people misjudge how long they lay awake in both directions. But sleep diaries remain the standard clinical intake instrument for insomnia, and CBT-I is built on them, which tells you something about what clinicians trust when the stakes are real. Actionability: highest, because a diary records the inputs you control. Attention cost: about ninety seconds a day.
| Tracks the lab standard | Actionable | Attention cost | |
|---|---|---|---|
| AI Coach narrative | Weakest — inference on inference | Generic | High |
| Raw metrics, weekly trend | Decent for sleep/wake, poor for stages | Moderate | Moderate |
| Diary + fixed wake time | Subjective, but clinically used | Highest | Low |
The verdict is not that the wearable is useless. Its total-sleep-time trend over three weeks is genuinely informative and no diary will match it for consistency. The verdict is narrower: the layer the update lets you hide is the layer that adds the least and costs the most. You are turning off the interpretation, not the instrument.
There is a documented failure mode here. Baron et al. (2017) published a case series in the Journal of Clinical Sleep Medicine coining the term orthosomnia — patients whose sleep worsened as they pursued perfect tracker scores, some of whom trusted device output over their own experience and over their clinician's reading. It was a handful of cases, not an epidemiological finding, and it gets cited far past what a case series can support. Call it plausible but thin. The mechanism it proposes — that anxious attention to sleep prolongs sleep onset — is not thin at all; that one is well-established, and it predates wearables by decades.
But does the AI Coach actually know what I drank?
No. It has no caffeine data unless you log it manually, and almost nobody logs it manually. When a coaching summary connects your poor night to "lifestyle factors," it is pattern-matching on heart rate and movement, not on your 4 p.m. cortado.
Which is a shame, because caffeine timing is one of the few sleep variables with clean experimental evidence behind it. Drake et al. (2013), in the Journal of Clinical Sleep Medicine, gave 12 subjects 400 mg of caffeine — roughly a large drip coffee, though bean and brew make that range wide — at 0, 3, and 6 hours before bed, with placebo controls and in-home polysomnography. The six-hours-before dose still measurably reduced total sleep time, by more than an hour in the aggregate measure. Twelve people is a small study and the dose is on the high side of typical. But the direction has held up across replications, and it is grounded in pharmacokinetics rather than correlation: caffeine's half-life in healthy adults runs about 4 to 6 hours, meaning a 200 mg afternoon dose still has roughly 100 mg circulating at bedtime, competing with adenosine at A1 and A2A receptors — blocking the accumulated sleep-pressure signal your brain spent all day building.
That is a mechanism you can act on. "Your readiness is low, consider prioritizing recovery" is not.
An honest rule of thumb
Trust your tracker for duration and trend. Do not trust it for stages. And treat any sentence it generates about why your sleep was bad as a hypothesis it has no way to test.
If you want a number to act on tonight: count back 8 hours from your intended bedtime and make that your last caffeine of the day. Not because 8 is magic — it is roughly one and a half to two half-lives, enough to drop a normal afternoon dose to a quarter of its peak. If you are a fast metabolizer, you will get away with less. Most people don't know which they are, and the cheapest way to find out is to test it rather than to read about it.
What to try this week
Hide the coach card for seven nights. Keep exactly two things: a fixed wake time, alarm set, weekends included — and a one-line note each night of when you had your last caffeinated drink and what it was. Nothing else. At the end of the week, look at the app's total-sleep-time trend, which was collecting quietly the whole time, and lay it next to your seven notes.
You will have done, at small scale, the thing the coach cannot: connected a variable you controlled to an outcome you measured.
The most useful feature in a health app is often the one that agrees to stop talking.