Automated sleep stage and event detection algorithms using quality-controlled polysomnography annotations.
Study Objectives To develop machine learning models for sleep stage classification, arousal detection, and respiratory event detection from overnight polysomnography, and to evaluate their performance relative to expert scorers. Methods Overnight polysomnography recordings were obtained from healthy...
| Publicado en: | Sleep Advances Vol. 7; no. 2; pp. 1 - 18 |
|---|---|
| Autores principales: | , , , , , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Oxford University Press / USA
2026
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=195378601&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 195378601 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 26325012 MB4X jtl: Sleep Advances issn: 26325012 maglogo: N pubinfo: dt: 2026 vid: 7 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 195378601 195378601 195378601 10.1093/sleepadvances/zpag054 195378601 ppf: 1 ppct: 17 formats: tig: atl: Automated sleep stage and event detection algorithms using quality-controlled polysomnography annotations. aug: au: Kaneda, Michiru Ogaki, Sho Nohara, Tomoyuki Fujita, Syuhei Osako, Naoshi Yagi, Tomoko Tomita, Yasuhiro Ogata, Takanori affil: ACCELStars, Inc., Tokyo, Japan sug: subj: Automation Sleep Stages Classification Arousal Evaluation Detection Algorithms Polysomnography Machine Learning Algorithms Classification Algorithms Sleep Apnea Syndromes Diagnosis Sensitivity and Specificity Evaluation Prediction Models Diagnosis, Computer Assisted Human Male Female Adult Middle Age Aged Consensus Funding Source Comparative Studies Interrater Reliability Decision Trees kappa Statistic Questionnaires Sleep Apnea Syndromes Electroencephalography Electromyography Electrooculography Descriptive Statistics Machine Learning Decision Support Systems, Clinical Validation Studies Precision Sleep Apnea, Central Diagnosis Adult: 19-44 years Middle Aged: 45-64 years Aged: 65+ years Male Female ab: Study Objectives To develop machine learning models for sleep stage classification, arousal detection, and respiratory event detection from overnight polysomnography, and to evaluate their performance relative to expert scorers. Methods Overnight polysomnography recordings were obtained from healthy participants and participants referred for suspected sleep-disordered breathing. Four certified scorers completed calibration sessions and generated reference annotations for sleep stages, arousals, and respiratory events. A subset of recordings was independently annotated by all scorers to support consensus analyses, enabling direct comparison between model outputs and human inter-scorer agreement. Gradient-boosted decision tree models were trained using hand-crafted features derived from standard physiological signals. Results Sleep stage classification achieved an accuracy of 0.840, a Cohen's kappa of 0.791, and an F1-score of 0.841, with limits of agreement for total sleep time of approximately ±0.5 h. Arousal detection achieved an F1-score of 0.733, with limits of agreement for the arousal index of approximately ±15 events/h. Respiratory event detection achieved an F1-score of 0.818, with limits of agreement for the apnea–hypopnea index also within approximately ±15 events/h. In consensus analyses, model performance was comparable to human inter-scorer agreement for sleep stages and arousals, while remaining below human inter-scorer agreement for respiratory events, despite high absolute performance relative to prior studies. Conclusions The proposed models achieved performance approaching human-level agreement across major sleep scoring tasks. These findings indicate that high consistency in expert annotations is a key factor underlying robust model performance and support the use of quality-controlled annotations for developing reliable automated sleep analysis systems. Statement of Significance Manual scoring of overnight sleep studies remains a major bottleneck in sleep medicine, limiting efficiency, consistency, and large-scale research. This study demonstrates that interpretable automated analysis can achieve performance approaching human-level agreement for core sleep scoring tasks when reference annotations are highly consistent. By directly comparing model outputs with calibrated inter-scorer agreement, the results show that annotation quality is a key determinant of attainable accuracy, rather than model complexity alone. Such systems may provide stable and reproducible reference outputs that support clinical decision-making, scorer training, and standardization across centers. Important remaining challenges include validation across institutions and populations, robustness to real-world signal artifacts, and extension to clinically meaningful subtypes of respiratory events. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|