Every morning for six years, a small computer on my wrist has issued a verdict on the previous eight hours and compressed it into two digits. Last night: 74. Fair.

I do not know what a 74 is. Neither, in any externally auditable sense, does the company that assigned it.

That gap sits at the center of the cost-benefit case for Fitbit wearables — and for the whole consumer sleep-tracking category — and it got more interesting once the data finally started moving. For most of the last decade, getting numbers off a Fitbit and into Apple Health on an iPhone meant paying a stranger on the App Store four dollars for a bridge app and then hoping their API access survived the next quarter. Now the sync is first-party. The plumbing works.

Which retires one question and replaces it with a better one. Not can I consolidate this, but what is the consolidated thing worth?

What most people do

The pattern is consistent enough to be a genre. You buy a tracker somewhere between $100 and $250. You wear it to bed, which means you also charge it at some odd hour, usually while you shower. You look at the score in the first ninety seconds of being awake, before coffee, before your eyes have fully accommodated. If the number is high you feel briefly vindicated. If it is low you spend the morning reinterpreting your own alertness through it — a phenomenon that is very hard to unsee once you've noticed yourself doing it.

Then, if you're on an iPhone, you consolidate. Everything into Apple Health, one archive, one place the rings close. This is a reasonable instinct and I share it. Health data that lives in four proprietary apps is data you will never look at again after you switch devices.

The costs are easy to tally and mostly small. Hardware, amortized: maybe $60 a year. Fitbit Premium, if you take it: $9.99 a month or $79.99 a year, which buys you more granular charts and a score breakdown. Attention: thirty seconds a morning is about three hours a year, which sounds trivial until you notice it's three hours spent forming an opinion about your body based on a proprietary composite.

The cost nobody prices is the interesting one. Baron and colleagues (2017), writing in the Journal of Clinical Sleep Medicine, described a small case series of patients who arrived at a sleep clinic distressed by tracker data, sometimes with objectively adequate sleep. They named it orthosomnia — the pursuit of perfect sleep as measured, at the expense of sleep as slept. It's a case series, not a trial. It proves the failure mode exists, not how common it is. But if you have ever lain awake calculating how the current hour will render in tomorrow's chart, you already know the mechanism is real.

And here is what almost nobody does: log the inputs. The tracker measures an output — last night — with no channel for the variables that produced it. What time was your last coffee? How much? Was it the 180 mg cold brew or the 65 mg cortado? Without that column, the correlation you actually want cannot be computed by you, by the app, or by Apple Health. You have built a beautiful instrument that records the weather and never asks where the clouds came from.

Does Fitbit data sync to Apple Health on iPhone?

Yes. On current versions, the Fitbit app for iOS can write its core metrics directly into Apple Health — you enable it from the account or connected-apps area of the Fitbit app, then grant per-category permissions on the Apple side. No third-party bridge, no subscription to a middleman. (The exact menu location has moved more than once; if you don't see it, update the app before you go looking for a workaround.)

What crosses is the summary layer: sleep start and end times, total sleep duration, the stage breakdown as Fitbit's classifier assigned it, steps, resting heart rate. What does not cross is the raw signal — the accelerometer stream, the photoplethysmography waveform, the beat-to-beat intervals the stage classifier was actually built on. You are importing conclusions, not evidence. That distinction matters for everything below, and it is the one part of this arrangement worth being genuinely annoyed about, because the sensor is on your wrist and the interpretation is on someone's server.

What the evidence suggests

How a wrist tracker decides you're asleep, in the order it happens

A three-axis accelerometer samples wrist motion, typically in the tens of hertz. Those samples get collapsed into movement counts per 30-second epoch — the same unit clinical actigraphy has used since the 1980s. Meanwhile a green LED on the underside of the case fires into your skin, and a photodiode measures how much light comes back; blood absorbs green, so the reflected signal pulses with each heartbeat. From that waveform the device derives heart rate, beat-to-beat variability, and a respiratory-rate estimate from how breathing modulates the pulse amplitude.

A photorealistic environmental portrait of a person sitting on the edge of an unmade…

Then a proprietary classifier takes those three streams — motion, heart-rate variability, breathing — and labels each epoch: wake, light, deep, REM. That label is a statistical guess trained against polysomnography in a lab population. It is not a measurement of your brain, because nothing on your wrist can see your brain.

The consequences follow directly from the physics. Lie perfectly still, awake, with a calm heart rate, and you look exactly like a sleeping person. This is why validation studies keep finding the same asymmetry: high sensitivity, poor specificity. Chinoy et al. (2021), in Sleep, ran seven consumer devices against in-lab polysomnography in 34 healthy adults and found most were very good at calling sleep when sleep was happening and considerably worse at catching wake once you were in bed. de Zambotti et al. (2018), in Chronobiology International, compared a Fitbit Charge 2 against polysomnography in 35 subjects and reported sensitivity for sleep around 0.96 against specificity for wake nearer 0.6 — with stage agreement that was better than chance and well short of clinical.

So, sorted honestly:

  • Well-established: wrist devices track sleep timing and duration usefully, and overestimate both, typically by ten to twenty minutes.
  • Plausible but thin: overnight resting heart rate and HRV trends as a signal of load, illness, or alcohol. The direction is consistent; the effect size for any one person is not well characterized.
  • Folk wisdom: that your "deep sleep percentage" is a number you can train, and that a sleep score is a thing worth optimizing.
What Apple Health receives How it's actually derived Worth trusting for
Sleep start / end, total sleep Motion + heart-rate classifier Timing and consistency; runs 10–20 min long
Resting heart rate Overnight PPG, lowest sustained values Week-over-week trend — the best single number here
Sleep stages Classifier on HRV and respiration Rough shape only; poor epoch-level agreement
Time awake after falling asleep Motion-dominated Little — this is precisely the weak axis
Sleep score Undisclosed proprietary composite Nothing externally validated. It's an interface element

Now the part that argues for instrumentation. Drake et al. (2013), in the Journal of Clinical Sleep Medicine, gave 12 healthy sleepers 400 mg of caffeine at 0, 3, and 6 hours before bed.1 Even the six-hour dose cut measured sleep by more than an hour. The subjects did not reliably notice. That is the whole case for owning one of these things: a real effect on your sleep, from a decision you make in the afternoon, invisible to introspection — and it lands on total sleep time and sleep onset, which happen to be the two things a wrist device measures least badly.

What I actually do

The score notification is off. It has been for over a year. Apple Health is where the data lands and stays; it is an archive, not a dashboard, and I don't open it in the morning.

On Sunday I look at two numbers across the prior week. Median sleep midpoint — the clock time halfway between falling asleep and waking, which is the cleanest available proxy for circadian consistency. And median resting heart rate, because it's derived from a directly measured physical quantity rather than a classifier's opinion of one.

For caffeine I stopped monitoring and started testing. Passive observation across a life with too many variables produces nothing; four days at a 2 p.m. cutoff, then four days at 5 p.m., then repeat the pair, produces a comparison. Sixteen nights, two conditions, two outcome measures. It's a crude ABAB design and it's still infinitely more informative than staring at a score.

The honest rule of thumb, for tonight: if a number on your wrist can't change a decision you'd make before dinner, it isn't data, it's decor. Turn off the ones that fail that test. Two survive — your resting heart rate, and the clock time you actually fell asleep.

No Premium. The subscription buys resolution on measurements whose accuracy doesn't support that resolution.

The watch charges from 8 to 9 p.m. on the kitchen counter, next to the grinder, which is empty by then because my last coffee is at 11:30 in the morning — not because a chart told me so, but because the sixteen-night version of that experiment came back unambiguous and I stopped arguing with it. The score is still generated every morning; I just never see it. The one number I kept was never trying to grade me.


  1. 400 mg is roughly four cups of drip coffee taken at once, which is a large but not exotic dose. Caffeine's half-life runs about 4 to 6 hours in most adults, and CYP1A2 variation moves that considerably — which is the honest reason blanket cutoff times are always a little wrong for somebody.