ML-based screening pipeline enables 94% sensitive, 98% specific NT1 detection from brief questionnaire + genetic marker
Stanford Sleep Medicine Center, USA · March 2025

Narcolepsy Type 1 (NT1) is a chronic neurological disorder caused by hypocretin deficiency, characterized by excessive daytime sleepiness and cataplexy—sudden episodes of muscle tone loss triggered by emotions. Accurate diagnosis in large-scale settings remains challenging due to NT1's low prevalence (~30 per 100,000), reliance on invasive testing (polysomnography, MSLT, CSF hypocretin assays), and diagnostic delays averaging 5–10 years from symptom onset.
In sleep clinics, rapid risk stratification is critical to prioritize severe patients for confirmatory testing while minimizing unnecessary referrals. Current screening tools (ESS alone, rule-based decision trees) lack the sensitivity and specificity needed to scale NT1 detection across diverse populations.
We developed machine learning classifiers trained on the Stanford Cataplexy Questionnaire and optional HLA-DQB1*06:02 genotyping to enable high-specificity NT1 risk stratification in clinical and population-based settings. The approach prioritizes specificity to reduce false positives in low-prevalence populations, then implements a post-hoc veto rule to correct any predicted NT1 cases lacking the HLA biomarker, further improving diagnostic precision.
The framework was designed for immediate clinical adoption: it uses only self-reported symptom data (emotional triggers for cataplexy, specific muscle weakness patterns) plus an optional genetic test, eliminating dependence on costly MSLT or sleep studies for initial triage.
Full cataplexy questionnaire (k=27) without HLA achieved near-perfect discrimination:
Adding HLA-DQB1*06:02 (present in 91.8% of NT1 cases vs. 24.6% of controls) increased specificity to 99%+ with zero false positives in full feature set. The post-hoc veto rule further improved specificity without sacrificing sensitivity, reducing false positives from 81 (ESS-only model) to near-zero.
The 10-item reduced questionnaire (designed for scalability in large cohorts like UK Biobank) performed comparably:
| Metric | Full Feature Set (k=27) | Reduced Set (k=10) | |--------|------------------------|--------------------| | AUC | 0.9955 | 0.995 | | Sensitivity | 94.85% | 94.62% | | Specificity | 98.05% | 98.24% | | NPV | ≥0.99998 | ≥0.99998 | | Model specificity with HLA veto rule | 99%+ | 99%+ | | Training cohort | 1,207 (280 NT1) | 1,207 (280 NT1) |
This work demonstrates that a brief 10-item cataplexy questionnaire, analyzed via machine learning and optionally paired with HLA genotyping, can reliably identify high-risk NT1 patients for clinical confirmation. For sleep clinics, the approach enables:
Rapid risk stratification: Severe patients (exhibiting laughter-triggered cataplexy + focal muscle weakness + high ESS) are flagged for immediate MSLT priority, reducing diagnostic delays.
Diagnostic tool validation: Clinicians can assess questionnaire relevance in their patient population—comparing false positive rates, feature importance, and sensitivity at different thresholds.
Scalable screening: The reduced 10-item feature set integrates directly with large population cohorts (UK Biobank) and EHR systems, requiring no additional data collection beyond standard intake interviews.
False-positive reduction: HLA-DQB1*06:02 testing (if available) dramatically reduces unnecessary referrals while maintaining 94%+ sensitivity, making the tool practical for population-level use.
Future validation in independent cohorts, combined with wearable actigraphy or GWAS, may further improve positive predictive value for population-based screening initiatives.