Garmin CIRQA matched a reference deep-sleep reading 76.5 percent of the time, but its REM tracking was weak enough that it sometimes missed REM phases entirely. That split matters most for athletes, recovery obsessives, and Garmin loyalists who want sleep data they can actually act on, not just a clean-looking score.
The finding comes from a review by postdoc and YouTuber Rob ter Horst, published on The Quantified Scientist and covered by Notebookcheck. The short version: Garmin CIRQA looks unusually strong at identifying deep sleep in this initial test, but the evidence does not validate the whole sleep stack.
That is the real story beneath the headline. CIRQA may be a step forward for Garmin’s display-less wearable strategy, but the test also shows why consumer sleep tracking remains a precision trap. A wearable can look excellent on one sleep stage and unreliable on another.
For product context, the device sits in the same Garmin wearables conversation as Garmin CIRQA Leak Ditches Screens for the Wrist War, while Garmin’s broader software-driven device maintenance can be seen in Garmin Instinct Update Kills Two Bugs Owners Feel Daily.
Garmin CIRQA’s deep-sleep win gives buyers one metric to respect, not a full sleep verdict
The most impressive result is narrow but real: deep sleep detection aligned with the reference 76.5 percent of the time. Rob ter Horst described that as one of the best results he has seen, according to Notebookcheck.
That is not a small achievement. Deep sleep is one of the stages users tend to obsess over because it feeds recovery narratives across wearable apps. If CIRQA is reliably better at flagging those periods for this tester, Garmin has a credible strength to build around.
But the buyer question is sharper: does one strong sleep-stage result make the whole device trustworthy?
Not yet. REM sleep was the weak point. Notebookcheck reports that CIRQA not only miscalculated REM duration, but at times failed to detect REM phases at all. A supplied secondary summary of the same review reports 46 percent REM agreement and 64 percent light-sleep agreement, while rounding deep sleep to 77 percent.
That mix makes CIRQA useful in pieces. It does not make it a sleep-lab replacement, and it does not make every morning score equally credible.
The test protocol helps makers, but the reference system limits what can be proven
This was more meaningful than a casual “my watch said I slept badly” impression. The CIRQA was compared against a reference device, and Rob ter Horst’s testing style is structured around side-by-side measurement rather than vibes.
Still, the protocol has two big constraints.
First, Notebookcheck says the data is based on only one test subject. That matters because sleep physiology, skin tone, wrist fit, movement, and nighttime behavior can all change how a wearable performs. The source does not provide a multi-person sample, so the result cannot be generalized with confidence.
Second, the reference was a Hypnodyne ZMax, which primarily measures brain waves from the forehead. Notebookcheck flags a key limitation: during REM sleep, brain-wave patterns can resemble light sleep or even wakefulness because of desynchronized EEG signals.
That creates a measurement problem. If the reference itself struggles to distinguish REM cleanly, CIRQA may not be solely responsible for every mismatch.
Clinical sleep monitoring uses more than forehead EEG. Notebookcheck notes that it also uses electrodes near the eyes and on the chin. Those measure eye movements and muscle relaxation — the signals that help identify REM more directly.
The builder question is obvious: are Garmin’s algorithms failing, or is part of the benchmark ambiguous?
The answer is likely some of both, but the supplied evidence cannot split the blame cleanly.
The numbers show a wearable that tracks steady signals better than fast-changing ones
CIRQA’s performance looks strongest when the underlying physiology is stable. That pattern appears in both sleep and exercise data.
| Test area | Reported CIRQA result | MLXIO analysis |
|---|---|---|
| Deep sleep | 76.5 percent agreement with reference | Strongest reported sleep-stage result |
| REM sleep | Sometimes missed entirely; supplied summary reports 46 percent agreement | Main sleep-stage weakness |
| Light sleep | Supplied summary reports 64 percent agreement | Middling, and likely tangled with REM misclassification |
| Steady running | Good heart-rate results | Works better when heart rate changes gradually |
| Interval training | Overestimated versus Polar H10 in several instances | Problem is not just optical-sensor lag |
| Indoor cycling | Supplied summary reports 0.99 correlation | Strong controlled-exercise result |
| Outdoor biking | Supplied summary reports 0.94 correlation | Still high, but less controlled than indoor cycling |
The exercise results matter because they echo the sleep findings. CIRQA seems more dependable when signals are smooth and repetitive. Steady running looks good. Indoor cycling looks very strong. Nighttime heart-rate tracking also appears solid in the supplied summary.
Intervals expose the weakness. Notebookcheck says the main issue was not the usual optical-sensor lag, but overestimation: CIRQA recorded significantly higher values than the Polar H10 reference in several instances.
For typical users, Notebookcheck says that may not matter unless they are running hard intervals and trying to control intensity strictly by heart rate.
The athlete question: if CIRQA overreads during intensity spikes, can you trust it for training control?
For steady-state work, probably more often. For strict interval pacing by heart rate, the source material says caution is justified.
End users should treat CIRQA as a trend tool before treating it as a sleep judge
CIRQA’s best use case is not diagnosing sleep architecture. It is tracking patterns across ordinary nights.
A user with consistent bedtimes and stable routines may get value from repeated deep-sleep and heart-rate trends. A single night with an odd REM readout should carry less weight. The source supports that distinction because CIRQA’s performance is uneven across metrics, not uniformly poor or uniformly strong.
The practical split:
- Useful: multi-night trends in deep sleep, general nighttime heart rate, and steady-activity heart-rate patterns.
- Risky: treating REM duration as precise, especially when the device may miss REM phases.
- Most questionable: using single-night stage breakdowns to make aggressive training, recovery, or lifestyle decisions.
This is where “more data” can mislead. Wearables infer sleep from signals such as movement and cardiovascular patterns. Even when a device produces confident-looking graphs, the underlying physiology can be ambiguous.
The buyer question: should CIRQA change your behavior tomorrow morning?
Only if the signal repeats. One strange sleep-stage chart should be treated as noise until it becomes a pattern.
Garmin’s sleep-test result is neither breakthrough nor flop
The CIRQA result is more interesting than a simple pass/fail. Deep sleep performance looks genuinely notable in this initial test. REM performance does not.
That makes “sleep breakthrough” too generous and “precision flop” too harsh.
For Garmin, the opportunity is clear: tighten REM detection and reduce heart-rate overestimation during fast intensity changes. The source material does not prove whether that requires better algorithms, different sensor handling, or both. It does show where the pain points are.
For users, the right reading is conditional confidence. Trust CIRQA more when it tracks stable, repeatable patterns. Trust it less when it claims high precision during messy biological states: REM sleep, brief awakenings, hard intervals, and anything with rapid transitions.
For clinicians, the evidence remains far too thin. One test subject and a non-clinical reference setup cannot support medical-grade claims.
The next evidence to watch is specific: more subjects, more nights, clearer REM validation, and comparisons against fuller clinical monitoring that includes eye and chin measurements. If CIRQA keeps its deep-sleep strength while improving REM agreement, Garmin has a stronger sleep story. If REM misses persist, the device remains a promising recovery tracker with a blind spot in one of the sleep stages users care about most.
Key Takeaways
- CIRQA appears unusually strong at detecting deep sleep, a key metric for recovery-minded users.
- Weak REM tracking shows Garmin’s sleep data still should not be treated as fully precise.
- The test highlights the broader limits of consumer wearables that can excel at one metric while failing at another.









