There is a habit, common to anyone who has worn a fitness tracker for more than a few months, of treating the numbers on the screen as measurements. They are not measurements in the strict sense. They are estimates produced by an algorithm reading a noisy optical signal through skin, a three-axis accelerometer, and in some cases a thermistor or an electrode pair. Some of those estimates are very good. Some are honest approximations. A few are essentially decoration.
What follows is a sensor-by-sensor map of where consumer wearables actually hold up against the laboratory tools they are imitating, and where they fall apart. We are going to name studies, give numbers, and point at the failure modes you can reproduce yourself. The short version is at the bottom, but the short version is not the article.
Wrist optical heart rate (PPG)
This is the sensor every modern wearable revolves around. A green LED (usually around 525 nm), sometimes paired with red and infrared LEDs, shines into the skin. A photodiode catches what scatters back. The small periodic modulation in that signal tracks the pulse wave, and an algorithm turns that into a number on the screen.
At rest, the technique works well. Independent comparisons against a Polar H7 chest strap have repeatedly put wrist PPG within about 2-3 BPM at rest and during steady, low-intensity cardio. This is the regime PPG was designed for and the regime most casual users spend most of their time in. If your primary use case is glancing at your resting heart rate over coffee, the wrist sensor is doing its job.
The wheels come off when the wrist starts moving fast or the heart rate changes quickly. Bent et al. 2020, in npj Digital Medicine (“Investigating sources of inaccuracy in wearable optical heart rate sensors”), put six wrist devices through a battery of activities and identified four reliable accuracy killers: motion (especially wrist-heavy motion), low perfusion (cold hands), skin tone (more on this in a second), and the rate of change of heart rate itself. During HIIT and similar transitions, MAPE on most devices crossed 10%, with worst-case errors above 30 BPM on individual readings. The same paper found smaller but measurable differences across skin tones for some devices; motion and perfusion almost always dominated, but pigmentation was not zero.
Other things that degrade the optical signal in ways you can verify on your own wrist: a tattoo under the sensor (ink absorbs green light unevenly), a loose band, cold hands after a long walk in winter, sweat pooling under the sensor on a humid run, and rowing strokes or kettlebell swings that whip the wrist through fast accelerations. None of this means the sensor is broken. It means PPG has a regime where it is excellent and a regime where it is unreliable, and consumer marketing tends to blur the two. We collected the cross-device numbers in a separate piece on fitness tracker accuracy.
The history of this is part of why we keep writing about the Basis Peak, which in 2014 was one of the first wrist trackers to attempt continuous PPG and ran into nearly every one of these failure modes years before they had names.
HRV measured at the wrist
Heart-rate variability, the beat-to-beat variation in the gap between successive R-peaks, is the most useful single number most wearables produce overnight. It correlates with autonomic state, training load, sleep quality, and (less reliably) stress.
The catch: chest straps measure HRV from the ECG, which gives you precise R-peak timing. Wrist PPG infers it from the optical pulse wave, which is noisier and which arrives at the wrist a few hundred milliseconds after the actual heartbeat. Algorithms can recover most of the meaningful HRV signal from a clean overnight PPG trace, especially during deep sleep when the wrist is still, but the absolute numbers do not directly compare to a chest-strap HRV reading, and they do not compare across devices either.
This is the part of the data that gets misread most often. Your Apple Watch HRV and your friend’s Whoop HRV are computed differently, sampled differently, and averaged over different windows. Comparing them is meaningless. Comparing your HRV to your own HRV from last week is the real use case. We wrote a longer piece on what HRV actually is that goes into the autonomic-nervous-system side; the takeaway is that wrist HRV is a trend instrument, not a calibrated one.
SpO2 from the wrist
This is the sensor where the gap between marketing and measurement is widest.
Pulse oximetry on a finger works because the fingertip is thin, well-perfused, and easy to clamp between an LED and a photodiode on opposite sides of the tissue (transmissive mode). The wrist is none of those things. Wrist SpO2 has to work in reflectance mode, with the LED and the photodiode on the same surface, looking at a thicker, less-perfused, more motion-prone piece of anatomy. The signal is fundamentally weaker.
In practice, wrist SpO2 drifts 2-4 percentage points from a finger oximeter under the same conditions, in most peer-reviewed comparisons. That gap is fine for noticing a long-term trend or spotting a possible apnea pattern over weeks. It is not fine for any decision that hinges on whether your saturation is 92% or 95%. Of the major consumer wrist trackers, only a handful have any FDA-cleared SpO2 functionality, and even those are usually cleared for “wellness,” not for clinical oximetry. We have a deeper write-up at our SpO2 wearables explainer.
The dropouts are also worth knowing. Wrist SpO2 fails silently when the band is loose, when the wrist is held below the heart for too long, when the wrist is cold, and when the wearer has a tattoo or significant arm hair under the sensor. The short version is the same as for everything else on this list: trend, not number.
Skin temperature
The skin-temperature sensors in modern wearables (Apple Watch, Oura, Fitbit Sense, Whoop) measure peripheral skin temperature, not core body temperature. Those are different physiological signals. Skin temperature reflects ambient conditions, blood flow to the periphery, the time of night, and a slow circadian rhythm. Core temperature reflects metabolic state and is what a fever screen actually needs.
This is why every wearable that ships a skin-temperature sensor describes it as a relative measurement and reports the nightly value as a deviation from a baseline rather than an absolute number. The deviations are useful. A persistent +0.5 °C overnight skin-temperature elevation, sustained over multiple nights, lines up reasonably well with the start of the luteal phase of the menstrual cycle, with the prodromal phase of a respiratory infection, and (less reliably) with a hangover.
What you cannot do with the skin-temperature sensor on your watch is screen for fever. That is not a limitation of any particular product. It is a property of the measurement site. Our piece on wearable skin-temperature sensors goes into the physiology in more detail.
Sleep staging
Polysomnography is the reference test for sleep architecture. It uses EEG to read brain activity, EOG to read eye movements, EMG for muscle tone, and ECG for heart rhythm. A consumer wearable has none of those. It has wrist movement and PPG.
That a wrist-worn device can do anything useful with sleep staging from those two inputs is mildly surprising, and it can. Total sleep time, the simplest output, is usually within 15-30 minutes of PSG across the major consumer devices. Wake after sleep onset is also reasonable. Sleep efficiency, the ratio between the two, is therefore also reasonable.
The stages are where it falls apart. Chinoy et al. 2021 in Sleep (“Performance of seven consumer sleep-tracking devices compared with polysomnography”) compared seven consumer devices to PSG across two nights in 34 adults. Total sleep time was within roughly an hour of PSG for most devices. Stage-level agreement was a different story: REM and deep sleep classification accuracy was poor across the board, with epoch-by-epoch agreement for those stages well below what would be acceptable in a clinical context. Light sleep was usually overestimated because every algorithm defaults to “light sleep” when it cannot tell what is happening.
The practical implication is that the staged sleep graph on your watch is best read as a rough estimate of when you were asleep and roughly how restless you were, not as a precise breakdown of your REM and deep cycles. The total-sleep number is the trustworthy part. See our sleep stages explainer for what each stage is, and what you can actually do with the information.
Step counting
Steps are the oldest and best-validated wearable measurement. The basic accelerometer-plus-threshold algorithm has been around since the 1960s in mechanical pedometers, and modern devices have refined it considerably.
On normal walking, every major wearable is within about 5% of a true count. That is true for the Apple Watch, the Fitbit lineup, every recent Garmin, the Oura Ring (which is a finger device but uses the same physics), and even cheap fitness bands. Where step counts fall apart is the activities they were never designed to measure: cycling registers as a small step count even though no steps happened, pushing a stroller or a shopping cart often undercounts because the wrist is not swinging, treadmill walking on a steep incline tends to undercount slightly, and swimming usually counts zero or near zero unless the watch is in a dedicated swim mode.
If your daily step number jumps by 4,000 the day you mow the lawn, that is real signal: the wrist was swinging through accelerations the algorithm interpreted as steps, and many of them probably were. If you do a long indoor cycling session and the daily step number barely moves, that is also real signal, just inconvenient.
Calorie estimates
This is the wearable number we trust least, and we are not alone in that. There is no consumer wearable that directly measures energy expenditure. Every calorie figure on every watch is a model that takes heart rate, motion, your sex, your age, your weight, and your height, and produces an estimate of metabolic equivalent (MET) per minute, which it converts into kilocalories.
Shcherbina et al. 2017 in the Journal of Personalized Medicine (“Accuracy in wrist-worn, sensor-based measurements of heart rate and energy expenditure in a diverse cohort”) put seven wrist devices through indirect calorimetry, the actual gold standard for energy expenditure, in 60 volunteers across sitting, walking, running, and cycling. The result for heart rate was reasonable: a median error around 5% under most conditions. The result for energy expenditure was not: median errors ran from roughly 27% on the best device to 93% on the worst, and the direction of the error varied across activities and across individuals.
The takeaway is that the calorie number is not a measurement. It is a rough internal index. If your watch tells you that today’s run was 600 kcal and yesterday’s was 450 kcal, the relative ordering is probably right. The absolute numbers should not be treated as a deficit you can subtract from a daily food intake. Anyone who has tried to lose weight on the basis of wrist calorie data has run this experiment for themselves; the modal result is frustration.
Sensor accuracy at a glance
| Feature | Typical accuracy vs gold standard | Primary sources of error |
|---|---|---|
| Wrist HR at rest (Apple Watch, Garmin, Fitbit, Whoop vs Polar H7 chest strap) | Within 2-3 BPM in published comparisons | Loose band, tattoo under sensor, cold wrist, dark skin tone in some devices |
| Wrist HR during HIIT / rowing / kettlebell (same devices vs chest strap) | 5-10% MAPE typical, worst-case >20 BPM (Bent 2020) | Wrist motion, fast HR change, sweat, perfusion |
| Wrist HRV overnight (PPG-derived vs ECG HRV) | Useful for trend only, no direct comparison to chest-strap HRV | Algorithm differences across vendors, window length, motion artifact |
| Wrist SpO2 (Apple Watch Series 6-9, Fitbit Sense 2, Garmin vs finger pulse oximeter) | 2-4 percentage point drift, no FDA clearance for clinical oximetry on most devices | Reflectance geometry, perfusion, band tightness, motion |
| Skin temperature (Apple Watch, Oura, Fitbit Sense vs clinical thermometer) | Relative only, baseline-deviation reading | Ambient temperature, blood flow, time of night; not a fever screen |
| Total sleep time (consumer devices vs PSG, Chinoy 2021) | Usually within 15-30 minutes of PSG | Stillness misclassified as sleep, brief wake epochs missed |
| Sleep staging (same devices vs PSG) | Poor stage-level agreement, especially REM and deep | No EEG/EOG/EMG input, light-sleep default bias |
| Step count on walking (Apple Watch, Fitbit, Garmin vs hand count) | Within ~5% on flat walking | Cycling, pushing a cart, treadmill incline, swimming |
| Calorie burn (wrist devices vs indirect calorimetry, Shcherbina 2017) | 27-93% median error depending on device | No direct measurement; demographic + HR + motion model |
Related accuracy comparisons
These newer guides apply the same sensor-accuracy caveats to specific buying decisions:
- Why your wearable sleep score can be wrong - sleep-score formulas, quiet wakefulness, and stage-estimate limits.
- Smart ring vs Apple Watch for sleep and recovery - finger PPG against wrist PPG in practical use.
- Fitbit Charge 6 vs Garmin Venu 3 - everyday health tracking when both devices rely on wrist optical sensors.
- Apple Watch SE 3 vs Fitbit Charge 6 - sleep, ECG, and heart-health caveats in the budget tier.
What this means for how you use a wearable
The honest framing is that a consumer wearable is a self-comparison instrument. It tells you whether today is different from yesterday, this week is different from last week, this month is different from the same month last year. The trend lines on your own data are real, and they are the part of the product you should pay attention to.
What it is not is an absolute clinical instrument. Comparing your Apple Watch HRV to your friend’s Whoop HRV is meaningless. Comparing your SpO2 reading from your wrist to a fingertip oximeter is roughly meaningful but not equivalent. Comparing your watch’s calorie estimate to a calorimetry-derived number is not meaningful at all. Comparing your watch’s sleep staging to a sleep lab result is meaningful for total time and unhelpful for stages.
The ECG features on some watches are a partial exception, because they are actually FDA-cleared as Class II devices for the specific use of detecting atrial fibrillation in adults 22 and older. That is a narrow, validated clinical claim, and the studies behind it are solid. We wrote a separate piece on the ECG features on consumer smartwatches, because the rules for those are genuinely different from the rules for everything else on the same device.
For everything else, the rule of thumb is the same one a good clinician would apply to any home monitoring data: track trends, watch for sustained changes, and bring the data to a professional rather than acting on a single reading. The wearable on your wrist is a useful instrument operated by an amateur on themselves with no calibration step. That is fine for what it is. It is not what the marketing implies.