Three.
That is the number sitting underneath nearly everything you have read about smartwatch anxiety. Not three hundred. Not three thousand. Three patients, written up in a 2017 case series in the Journal of Clinical Sleep Medicine by Kelly Glazer Baron and colleagues, under the title "Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?" Three case reports from a sleep clinic. No control group. No comparison arm. No denominator. The paper coined a word — orthosomnia, the perfectionist pursuit of ideal sleep data — and the word got out.
It is now in health columns and podcast episodes and the intake conversations of sleep clinicians, and occasionally in the mouths of people diagnosing themselves at dinner. I have used it myself. It has the grammar of a diagnosis. It is not one; it appears in no diagnostic manual.1 It is a name three clinicians gave to something they kept seeing in their waiting room.
I am not telling you this to debunk it. I am telling you because I am, fairly obviously, the kind of person the word was invented for, and I would like to know what is actually known about people like me — as opposed to what is merely repeated about us.
Here are my credentials. In February, at 11:48 p.m., I walked twelve laps of a hotel parking structure in a city I had landed in that afternoon, in dress shoes, in the cold, because I was forty-one active calories short of a Move goal I set in 2019 and had not missed since. I knew while I was doing it that those calories were an estimate produced by an accelerometer and a green LED. I knew the number was not a fact about my body. I walked the laps.
What the three patients actually did
The case reports are short, and the details are the part that gets lost in translation, so they are worth having.
All three arrived at a sleep clinic with complaints about their sleep, and all three arrived holding data. Not vague impressions — printouts, screenshots, weeks of nightly percentages. In each case the clinicians noticed the same pattern: the patient trusted the device more than the clinic, more than their own reported experience, and in at least one instance more than the clinic's own overnight measurement, which disagreed with the wearable and was therefore treated as the thing that must be wrong.
The most clinically interesting behavior was the simplest. Patients were spending more time in bed in order to raise the number. That sounds reasonable and is precisely backward. Insomnia is maintained, in large part, by extending time in bed to catch sleep that isn't coming; the frontline treatment, cognitive behavioral therapy for insomnia, does the opposite — it compresses time in bed to consolidate sleep and rebuild the association between the bed and unconsciousness. So the device's implicit instruction and the actual treatment pointed in opposite directions, and the device had a chart.
What the authors claimed was modest. They proposed a name, described what they had seen, and said it deserved attention. They did not claim an epidemic. Everything epidemic-shaped that has been written since is downstream of readers, not of the paper.
What a case series can't tell you
A case series is a report of things that happened. It is the lowest tier of clinical evidence and it is not trying to be anything else. Its job is to raise a hand.
So: three patients, selected because they walked into a sleep clinic. That is sampling on the outcome. Estimating the frequency of smartwatch anxiety from that sample is like estimating how often coffee causes insomnia by interviewing people in an insomnia clinic who happen to be holding coffee. You will find a striking rate. It will mean nothing about the population.
What the paper cannot tell you, and what nine years of citation has not supplied: how many wearable users develop this pattern. Whether tracking produces the anxiety or anxious people simply track more and harder. Whether it fades on its own, the way most new-toy behaviors do, or entrenches. Whether the people it happens to were going to find something else to be exacting about.
So let me sort the claims by how much weight they'll hold:
Well-established. The design techniques these devices use work. Streaks, progress geometry, and prompt timing are not speculative — they come out of a behavioral literature that is decades deep and reliably replicated in commercial settings, which is why every app has them.
Plausible but thin. That wearables cause clinically meaningful anxiety in a meaningful share of users. The evidence here is case reports, cross-sectional surveys that ask people to self-report both the anxiety and the tracking, and a scatter of small experiments. Cross-sectional surveys cannot separate cause from selection, and everyone who publishes them says so in the limitations paragraph nobody quotes.
Folk wisdom. That "everyone" is stressed about their rings. Most people wear these things casually, ignore two-thirds of the notifications, and stop caring by March.
Can a smartwatch actually cause anxiety?
For most people, no — not in the clinical sense of an anxiety disorder. For a smaller group, yes, plausibly, in the ordinary sense: the device produces a stream of daily numbers that arrive looking like verdicts, and if you are already someone who converts numbers into verdicts, it hands you roughly a thousand more of them a year. The honest position is that we have good evidence the mechanism is real and no good evidence about how many people it catches. Anyone quoting you a prevalence figure for this is quoting a survey, not a measurement.
But there is one experiment that makes the mechanism concrete, and it is my favorite study in this whole area because it removes sleep from the equation entirely. Draganich and Erdal (2014), in the Journal of Experimental Psychology: Learning, Memory, and Cognition, told undergraduates — roughly fifty per experiment, so small — that a sensor had measured their REM sleep the previous night. Half were told they'd gotten about 16 percent REM, which they were informed was below average. Half were told about 29 percent, above average. The numbers were assigned at random and had nothing to do with anyone's actual sleep. Then everyone took a cognitive test.
The group told they'd slept well did better.
Nothing about their sleep differed. What differed was the number they had been handed that morning, and what they concluded from it. If a fabricated percentage can move performance on a task, a real-but-uncertain percentage delivered by a device you trust can move how your Tuesday feels. That is not mysticism. That is expectation doing what expectation does.
The number arrives without error bars
Here is the design decision that makes all of this worse, and it is not a scary one. It is just a rounding convention.
Chinoy and colleagues (2021), in Sleep, ran seven consumer sleep-tracking devices against polysomnography — the electrode-and-lab standard — in 34 healthy adults. The devices were reasonably good at one thing: detecting that sleep was happening. Sensitivity for sleep ran above 0.9 for most of them. They were poor at the complementary task: recognizing that you were awake. Specificity for wake was low across the board, in some cases well under half. In plain terms, a wrist tracker sees a still wrist and calls it sleep. Lying in the dark, furious, motionless, is a state these devices are structurally bad at distinguishing from stage 2.
Movement estimates fare worse. Shcherbina and colleagues (2017), in the Journal of Personalized Medicine, tested seven wrist devices in 60 participants and found heart rate was generally decent — median error under about 5 percent for most devices — while energy expenditure was not. The best device's median error was around 27 percent. The worst was north of 90.
None of this means the devices are useless. It means the number on the screen is an estimate with a confidence interval, and the interface has stripped the interval off. Your watch does not say "somewhere between 6 hours 10 and 7 hours 5, and we genuinely cannot tell restless wake from quiet sleep." It says 6:42. Two digits, no hedge, same typography as the time of day. The precision is a typographic choice, not a measurement property, and you respond to the typography.
How the loop closes, in the order it happens
Walk one day through it.
7:04 a.m. The overnight summary is waiting before you are conscious enough to evaluate it. B. J. Fogg's model from the Stanford Persuasive Technology Lab is the cleanest description of why this timing works: behavior happens when motivation, ability, and a prompt converge. At 7:04 the prompt is on the wrist already touching your face and the ability cost is zero. You have formed an opinion about your night before your feet hit the floor.
10:50 a.m. A stand reminder. Ten minutes of an hour left. The action costs nothing, the reward is immediate and certain, and the loop between prompt and payoff is under sixty seconds. This is the cheapest possible training trial, and you run about twelve of them a day.
4:00 p.m. You glance at the rings. This is where the geometry does work that a progress bar starting at zero would not. Nunes and Drèze (2006), in the Journal of Consumer Research, gave about 300 car-wash customers one of two loyalty cards: an eight-stamp card that arrived with two stamps already on it, or a ten-stamp card that arrived blank. Both required eight more washes. The pre-stamped card was completed at roughly double the rate — around a third of customers versus under a fifth. Artificial advancement increases effort. Your rings are never empty; by 4 p.m. you have accrued something, and abandoning it now feels like discarding rather than declining.
9:30 p.m. The streak. This is the strongest component and the one most often described wrongly. People say these devices work like slot machines — variable rewards, unpredictable payoffs. That is not what a ring is. A ring is a fixed schedule: do the thing, get the thing, every time, no surprise. The pull isn't randomness. It's an accumulating asset with a cliff at the end. Prospect theory (Tversky and Kahneman, 1992) put the median loss-aversion coefficient around 2.25 — losses weighted roughly twice as heavily as equivalent gains — and a 900-day streak is a large, precisely quantified thing that can only be lost, never gained again. It is worth noting that loss aversion's generality has taken real fire; Gal and Rucker (2018), in the Journal of Consumer Psychology, argued the effect is far more context-dependent than its textbook status implies. But whatever the coefficient, an object that took three years to build and can be destroyed by one flu is not symmetric.
11:48 p.m. The parking structure.
The part that runs underneath all of this — the checking, the small spike of uncertainty before the number loads, the arousal that follows — is where I have to be careful. Conditioned arousal is well described in insomnia: the bed becomes a cue for wakefulness through repeated pairing, and that is the basis of stimulus-control therapy. Extending that model to a nightly ritual of checking a screen that evaluates you is reasonable. It is also extrapolation. Nobody has instrumented it properly. I believe it happens to me. I would not present that as a finding.
Persuasive design ethics when nobody is the villain
The field has had an ethics literature since almost the beginning. Berdichevsky and Neuenschwander (1999), in Communications of the ACM, proposed what they called the golden rule of persuasion: don't use persuasive technology to induce a behavior you would not consent to being induced into yourself. It's a good rule and it is nearly unfalsifiable in practice, because the designer who built the streak almost certainly does like having a streak.
The real asymmetry is measurement, and it should feel familiar by now. A company can measure, cheaply and continuously, whether a feature increases engagement. It cannot easily measure whether the same feature makes two percent of users walk laps of a parking garage at midnight. The first number appears on a dashboard every morning. The second one has no owner, no instrument, and no denominator. Harm that isn't measured doesn't lose arguments; it doesn't get to be in the argument.
I want to be fair here, because a lot of writing about smartwatch anxiety is not. Apple does not sell your activity data to advertisers. There is no infinite feed. The rings are honest about wanting you to move, which is more than most software will admit about its goals. And in watchOS 11, in 2024, Apple shipped the ability to pause your rings — daily, weekly, monthly — and to set different goals for different days of the week. That is a rest day. It took nine years, and it is what it looks like when a design ethics question gets answered by shipping something rather than by publishing a manifesto about it.
The remaining question is defaults. Pausing requires knowing the setting exists and then choosing to spend the streak's protection on yourself, which is exactly the decision the streak is engineered to make expensive. A feature that only the already-liberated will find is a real improvement and an incomplete one.
What your watch is good at, and what it isn't
| Metric | How much to trust it | Use it for |
|---|---|---|
| Total sleep time | Moderate; typically off by tens of minutes | Rough weekly averages, not last night |
| Sleep stages | Low; this is the weakest output on the device | Nothing you'd change a decision over |
| Resting heart rate | Good as a relative trend | Spotting illness, alcohol, overtraining |
| Heart rate variability | Noisy night to night, meaningful over weeks | Multi-week direction only |
| Active calories | Poor; median errors of 27–90%+ across devices | Comparing you to you, never to a food label |
| Steps | Good while walking freely, poor when your hand is on a cart or stroller | Consistency, not precision |
An honest rule of thumb
One night is noise. Two weeks is signal. Do not change anything on the basis of a number that moved once; change something when a line moves and stays moved.
And the tell that matters isn't in the data at all. It's behavioral: if you have ever altered a real-world plan — left a game early, skipped a dinner, declined a nap you needed — to protect a metric, that is the thing to notice. Not the hours. The rearranging.
Try this, this week
One small experiment, seven days, paper and pen on the nightstand.
Each morning, before you look at your watch, write down two things. First, how rested you feel, one to ten, in your body, right then. Second, your guess at what the watch is going to say. Then look, and write down what it actually said. That's it. Thirty seconds a day.
At the end of the week you will have seven pairs, and they will tell you something a year of scrolling summaries cannot. If your rating and the number track each other, the watch is confirming what you already knew, which means you can trust your own read and check the device less. If they don't track — if you felt fine on the 5-hour night and wrecked on the 8 — then you have located the actual problem, and it isn't your sleep. It's that you have been letting an estimate overrule a measurement you were taking with your entire body.
Whichever way it comes out, keep this in mind while you look at the number: it never actually saw you sleep. It felt your wrist go still, made a guess, and rounded it to the minute.
-
The ortho- prefix is borrowed from orthorexia, the term for an unhealthy fixation on eating correctly, coined by Steven Bratman in 1997 in a yoga magazine. It is also not in any diagnostic manual, and it too is now used as though it were. Words describing the pathologies of self-improvement seem to travel much faster than the evidence for them, which is probably its own finding. ↩