On the morning of 14 July the app told me I'd slept badly because of "caffeine later in the day." My last coffee the previous day was at 09:14. I know this because I had weighed the beans — 18 grams, roughly 200 mg — and typed the time into a spreadsheet before drinking it, because I was three weeks into an experiment built entirely around that number. The watch had the correct data. The summary on top of it had written fiction.
So here is the verdict, stated once and plainly: the current crop of Google Health AI failures are not sensor failures, they are narration failures — the hardware measures something roughly real, and then a language model writes a confident story about it that nobody checked against the numbers sitting one screen away.
That distinction matters more than it sounds, and it's the reason I'd rather run this piece than another wrist-tracker accuracy shootout.
The myth: the AI layer is only a narrator
The thing I keep hearing from people who should know better — engineers, mostly, and a couple of very sharp health-data people — goes like this: the summary layer is cosmetic. The hard part is optical heart rate through skin, motion classification, the actigraphy. The AI on top is just reading a table you could read yourself. Worst case it's redundant. It can't be wrong about your night, because it's looking at your night.
It's a tidy argument and it survives right up until you check it. The premise smuggles in an assumption: that a system summarizing data can only lose information, never add it. Summarizers don't work that way. A model asked to explain a night's sleep in two friendly sentences will produce two friendly sentences whether or not the data supports two friendly sentences, and the sentences it produces are generated, not looked up. Nothing in the pipeline holds them to the table.
I wanted to know how often the gap between the table and the sentence actually opened, and how wide.
What I actually did
Twenty-one consecutive nights. Every coffee weighed on a 0.1 g scale and converted at 11 mg caffeine per gram of roasted bean — imprecise, but consistently imprecise, which is what matters for a within-subject comparison. Time of last dose recorded to the minute. Decaf logged separately at an assumed 8 mg per cup. Lights-out marked with a physical button press on a bedside remote so I had a timestamp that owed nothing to the watch's own guess.
Each morning I did two things before reading anything: exported the night's raw-ish data (stage durations, resting heart rate, HRV, wake events, the tracker's own onset estimate), and screenshotted the AI-written summary. Then I compared them. On nine of the nights I also wore a second, older tracker on the other wrist — not as ground truth, because it isn't, but as a second opinion.
My caffeine intake ranged from 0 mg to 412 mg per day, with the last dose falling between 07:40 and 16:05. Median bedtime 23:20. I am one person with one sleep pattern and one set of arteries, and none of what follows is a study.
The evidence
Five nights where the two artifacts disagreed. These are representative, not cherry-picked outliers — the disagreement rate across the full 21 was 8 nights, or 38%.
| Night | What I logged | What the exported data showed | What the summary said |
|---|---|---|---|
| 14 Jul | Last caffeine 09:14, 200 mg | Onset 22 min, 3 wake events | "Caffeine later in the day may have affected your rest" |
| 17 Jul | 0 mg all day | Onset 14 min, RHR 51 | "Your late workout likely delayed sleep onset" (no workout logged) |
| 19 Jul | 412 mg, last dose 16:05 | Onset 48 min, 6 wake events | "A restful night — keep it up" |
| 23 Jul | 180 mg, last dose 08:20 | 41 min gap flagged low-confidence | "You spent 41 minutes in deep sleep before waking" |
| 02 Aug | 200 mg, last dose 11:00 | Watch off wrist 01:10–03:40 | "Light sleep throughout the early morning" |
The 19 July line is the one that should bother you. That was my worst night of the three weeks by every measure the device itself recorded — longest onset, most fragmentation, highest overnight resting heart rate — and it followed the largest caffeine load, with a dose at 16:05 that any sleep clinician would flag on sight. The summary congratulated me.
The 23 July and 02 August lines are a different species of wrong, and worse. Both are gaps: one where the device marked its own confidence as low, one where the watch was physically off my wrist charging. In both cases the underlying record said we don't know, and the sentence on top said something specific and calm. Nothing in the interface told me the difference between a measured 41 minutes and an imagined one.
Mechanism: why the sentence drifts when the numbers don't
Four things are going on, and only one of them is the model being dumb.
Confidence is stripped at the boundary. The sensor pipeline knows perfectly well that it has a 2.5-hour hole in the record. That uncertainty exists as a field somewhere. It does not survive the trip into prose, because prose has no syntax for error bars unless you build one — and "light sleep throughout the early morning, probably, we had the watch off for a bit" is not a sentence that ships in a consumer wellness app.
Fluency is the product, so gaps get filled. A summarizer's job is a complete, readable paragraph. A gap in the input doesn't produce a gap in the output; it produces the most plausible continuation. Plausible continuations of "asleep at 01:10" and "awake at 06:20" is light sleep in between. That's not hallucination in the exotic sense. It's the system doing exactly what it was built to do, applied to a domain where the honest answer was silence.
Correlational engines are made to speak causally. "Caffeine later in the day may have affected your rest" is a causal claim dressed in a hedge. The model has an extremely strong prior linking bad sleep to late caffeine — it's true in the aggregate, it's all over the training data — and on 14 July it reached for that prior instead of for my actual log, which said 09:14. The hedge word "may" does no work here; the sentence still points at a cause that didn't exist. When I fed it a genuinely caffeine-wrecked night on 19 July, the prior didn't fire, because the night's summary statistics happened to land in a range the model reads as unremarkable.
n=1 pattern-matching is structurally unreliable and gets presented anyway. Three weeks of my nights is not enough to establish that anything causes anything about my sleep. The app doesn't know that, or knows it and says the thing regardless.
None of these are hard problems in the sense of being unsolved. They're product decisions. Health-AI failures of this kind get shipped because a confident paragraph tests better than an honest table.
What it gets right
Credit where it's earned, because a review where everything is bad is as useless as one where everything is good. The resting heart rate trend across 21 nights tracked my second device within 2 bpm on every single night — that's genuinely good hardware. Wake-event counts were plausible and internally consistent. The onset estimates, checked against my button press, ran a median of 6 minutes optimistic, which is respectable for a wrist optical sensor and better than I expected. Long-run trend lines, where individual-night error washes out, were the most useful thing on the device by a distance.
The layer I'd remove is the one doing the talking.
Who this is for, and who it isn't
Buy in anyway if: you read the charts and ignore the prose. If your instinct is to tap through to raw stage durations and RHR trends, the hardware earns its place on your wrist and the summary is a skippable ad for itself. Same if you're tracking a slow variable — a training block, a taper off caffeine over months — where night-to-night noise doesn't matter.
Don't, or at least don't trust the summaries, if: you're using sleep data to make a decision that matters. Managing a diagnosed sleep disorder, adjusting a medication with your doctor, deciding whether the 4 p.m. espresso is the thing wrecking you. The failure mode isn't "slightly off." It's a fluent, specific, unfalsifiable-looking sentence about a night the device didn't observe. Also don't, if you're the kind of person who'll be quietly demoralized by a machine telling you that you slept badly — I noticed on three separate mornings that I felt worse after reading the summary than before, and I'm the guy who knew it was making things up.
If there's a clear winner in this comparison, it's the export button.
The honest takeaway
I can't tell you whether these summaries are model-generated end to end, how much human review sits behind them, or whether the next update quietly fixes it — I asked, and I don't have an answer. What I can tell you is what 21 mornings of side-by-side comparison showed: the numbers were mostly fine, the narration was wrong 38% of the time, and it was most confidently wrong precisely where the data was thinnest.
Tonight, if the app writes you a sentence about your sleep, go find the number it came from before you believe it — and if there is no number, there was no night.