Every few months a respiratory device manufacturer announces a mask line, and the announcement carries a percentage. Satisfaction, usually — some large fraction of patients in an internal survey who found the interface comfortable. That number is not worthless. It is simply not the number that decides whether CPAP masks stay on a face at three in the morning, or whether the supplier gets paid in month four.
In the United States, that number is four.
Four hours a night, on 70 percent of nights, across 30 consecutive days somewhere in the first 90. That is the Medicare adherence standard for continued positive airway pressure coverage, and most commercial payers have copied it close to verbatim. A mask that delivers a patient to 3.9 hours is an administrative failure. A mask that delivers the same patient to 4.1 hours is a success story that will end up in a slide deck. The clinical distance between those two people is twelve minutes.
Where four hours came from
The threshold has a paper trail, though it is thinner than the authority with which it gets cited. The common citation is Kribbs et al. (1993), American Review of Respiratory Disease, which fitted covert monitors to 35 patients on nasal CPAP and recorded mask-on-at-pressure time rather than machine-on time. Two findings mattered. Patients overestimated their own use by roughly an hour a night. And when the authors needed a way to describe the distribution, they split the cohort into "regular users" at four or more hours on 70 percent of nights. Fewer than half the group cleared it.
Notice what that study was doing. It was characterizing behavior, not locating a dose. The cut point was chosen to make a spread of numbers legible, not because four hours was where any outcome curve bent. Somewhere between 1993 and the coverage determinations that followed, a descriptive convenience became a reimbursement gate, then a clinical shorthand, and it now gets repeated at conferences as though it were a physiologic constant. The lineage from paper to policy is reconstructed rather than documented. Nobody publishes the meeting minutes.
What four hours does not measure
The best available dose-response work is Weaver et al. (2007), Sleep, which tracked 149 patients and asked a simple question: how many hours of use does it take before a treated patient looks like an untreated normal person?
The answer depended entirely on what you measured. Subjective sleepiness on the Epworth normalized at around four hours. Objective alertness on the Maintenance of Wakefulness Test took closer to six. Functional status on the FOSQ took about seven and a half. The relationship was continuous — more hours, more benefit, no plateau where the curve went flat and the clinician could relax.
Read the reimbursement threshold against that. It sits exactly where a patient stops reporting sleepiness and nowhere near where they stop being impaired. A patient at four hours flat is compliant on paper and undertreated in the two domains that govern whether they can drive a truck, hold a shift, or remember a conversation.
This has consequences at trial scale. SAVE (McEvoy et al., 2016, New England Journal of Medicine) randomized 2,717 patients with OSA and established cardiovascular disease and found no reduction in cardiovascular events. Mean nightly adherence was 3.3 hours. The result has been argued over ever since, and the most defensible reading is not that CPAP fails to prevent strokes. It is that the question was asked at roughly half the dose that normalizes daytime function. The rate-limiting component of that trial was the interface.
Does mask type actually change how well the therapy works?
Yes — and mostly not through comfort. Mask selection changes the route of breathing, and route changes the pressure required to hold the pharynx open. Oronasal masks generally need higher pressures to achieve what nasal masks achieve, and tend to leave more residual respiratory events behind at whatever pressure they are given. This is a therapeutic difference, not a preference difference, and it does not show up anywhere in a satisfaction percentage.
The cleanest head-to-head is Rowland et al. (2018), Journal of Clinical Sleep Medicine: a randomized crossover in which roughly four dozen patients with moderate-to-severe OSA cycled through a nasal mask, nasal pillows, and an oronasal mask. Oronasal finished worst on residual AHI, drew the least nightly use, and was the least preferred. Nasal and pillows performed close to identically.
The mechanism is less settled than the direction. Andrade and colleagues, in work published across roughly 2016 to 2018, argue that oronasal interfaces act on the retroglossal airway — either through posterior mandibular displacement under strap tension, or through the oral flow route itself pressing the tongue back. The imaging is suggestive and the samples are small. Call the effect well-established and the anatomy plausible but thin.
What follows practically is uncomfortable for the field's default reflex. A patient leaks through the mouth, so the interface is escalated to full-face. That buys a seal and spends pressure.
How a mask fails, in the order it fails
Sleep onset. Pharyngeal dilator tone drops and the mandible rotates down and back by a few millimeters. That alone changes seal geometry, because cushions are fitted on an awake, upright, muscularly toned face, and the face they sit on at 2 a.m. is a different shape.
The cushion lip breaks contact — usually at the nasal bridge or along the lateral cheek where the strap vector is weakest. Unintentional leak climbs.
The blower answers with more flow, because it is a pressure-targeted device and that is the only move it has. Pressure at the mask is roughly held. Pressure at the airway may not be. More importantly, the flow signal the machine uses to detect apneas and hypopneas degrades under exactly these conditions, so on most auto-titrating platforms high leak suppresses or invalidates event scoring. The device becomes least trustworthy at the precise moment it is most needed.
If the leak takes an oral route, dry air crosses the mucosa for the next twenty minutes. The patient surfaces with a mouth like a paper bag. By then the impulse is not to reseat the cushion. It is to take the thing off.
Hours logged: 3.4. Payer verdict: noncompliant. Actual finding: a cushion that lost its seal about an hour after muscle tone dropped.
Three interface classes, honestly summarized
| Interface | What the evidence supports | Characteristic failure | Download signature |
|---|---|---|---|
| Nasal mask | Best-studied; lowest effective pressure | Oral venting when nasal resistance rises | Low baseline leak with sharp spikes |
| Nasal pillows | Near-equivalent to nasal in crossover data | Nares irritation above ~12 cm H₂O; dislodgement in side sleepers | Low leak, short nights |
| Oronasal | Justified in fixed nasal obstruction | Higher required pressure, higher residual AHI, more perimeter to seal | Elevated leak and elevated residual AHI together |
What the compliance download cannot see
Leak figures are not comparable across manufacturers. ResMed reports unintentional leak, having subtracted modeled intentional vent flow. Other platforms report total leak, vent included. The widely quoted 24 L/min flag is a manufacturer convention with a reasonable engineering rationale behind it, not a physiologic boundary, and using it to compare two masks across two device brands produces mostly noise.
Residual AHI is a related trap. What the device reports is a flow-derived estimate generated by a proprietary algorithm, not a scored index. It does not have EEG, so it cannot see arousals. It has limited ability to separate central from obstructive events without added signals. And its accuracy falls off under high leak — the same condition that produced the problem you are trying to diagnose.
So the two numbers most often used to adjudicate whether a mask is working are a vendor-specific flow estimate and a threshold invented to describe a 35-person cohort in 1993.
An honest rule of thumb
Before reordering an interface, put two things side by side: the hour the patient says they wake, and the 95th-percentile leak trace.
If leak climbs and use terminates within the same hour, it is a seal problem, and cushion size, material, and strap vector are the levers worth pulling. If use terminates while leak stays flat, the mask is not the reason — look at pressure intolerance, aerophagia, nasal congestion, comorbid insomnia, or the fact that the patient has a toddler. Swapping interfaces in that second case is expensive, and it costs the patient another two weeks of belief that the therapy is the problem.
One more: when a patient sits at 3.5 hours, resist the instinct to celebrate the jump to 4.1. Weaver's curve says the meaningful work is between four and seven.
The myth is that a mask works if the patient tolerates it, and that the compliance download tells you whether they did.
The more accurate version is that a mask works if it holds its seal through the hours when muscle tone is lowest and the mouth wants to open — and the compliance download tells you only whether the claim gets paid.