Cargando…

Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores

BACKGROUND. Polygenic risk scores (PRS) are linear combinations of genetic markers weighted by effect size that are commonly used to predict disease risk. For complex heritable diseases such as late onset Alzheimer’s disease (LOAD), PRS models fail to capture much of the heritability. Additionally,...

Descripción completa

Detalles Bibliográficos
Autores principales: Hermes, Stephen, Cady, Janet, Armentrout, Steven, O’Connor, James, Carlson, Sarah, Cruchaga, Carlos, Wingo, Thomas, Greytak, Ellen McRae
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Cold Spring Harbor Laboratory 2023
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9934790/
https://www.ncbi.nlm.nih.gov/pubmed/36798198
http://dx.doi.org/10.1101/2023.02.10.23285766
_version_ 1784889949771792384
author Hermes, Stephen
Cady, Janet
Armentrout, Steven
O’Connor, James
Carlson, Sarah
Cruchaga, Carlos
Wingo, Thomas
Greytak, Ellen McRae
author_facet Hermes, Stephen
Cady, Janet
Armentrout, Steven
O’Connor, James
Carlson, Sarah
Cruchaga, Carlos
Wingo, Thomas
Greytak, Ellen McRae
author_sort Hermes, Stephen
collection PubMed
description BACKGROUND. Polygenic risk scores (PRS) are linear combinations of genetic markers weighted by effect size that are commonly used to predict disease risk. For complex heritable diseases such as late onset Alzheimer’s disease (LOAD), PRS models fail to capture much of the heritability. Additionally, PRS models are highly dependent on the population structure of data on which effect sizes are assessed, and have poor generalizability to new data. OBJECTIVE. The goal of this study is to construct a paragenic risk score that, in addition to single genetic marker data used in PRS, incorporates epistatic interaction features and machine learning methods to predict lifetime risk for LOAD. METHODS. We construct a new state-of-the-art genetic model for lifetime risk of Alzheimer’s disease. Our approach innovates over PRS models in two ways: First, by directly incorporating epistatic interactions between SNP loci using an evolutionary algorithm guided by shared pathway information; and second, by estimating risk via an ensemble of machine learning models (gradient boosting machines and deep learning) instead of simple logistic regression. We compare the paragenic model to a PRS model from the literature trained on the same dataset. RESULTS. The paragenic model is significantly more accurate than the PRS model under 10-fold cross-validation, obtaining an AUC of 83% and near-clinically significant matched sensitivity/specificity of 75%, and remains significantly more accurate when evaluated on an independent holdout dataset. Additionally, the paragenic model maintains accuracy within APOE genotypes. CONCLUSION. Paragenic models show potential for improving lifetime disease risk prediction for complex heritable diseases such as LOAD over PRS models.
format Online
Article
Text
id pubmed-9934790
institution National Center for Biotechnology Information
language English
publishDate 2023
publisher Cold Spring Harbor Laboratory
record_format MEDLINE/PubMed
spelling pubmed-99347902023-02-17 Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores Hermes, Stephen Cady, Janet Armentrout, Steven O’Connor, James Carlson, Sarah Cruchaga, Carlos Wingo, Thomas Greytak, Ellen McRae medRxiv Article BACKGROUND. Polygenic risk scores (PRS) are linear combinations of genetic markers weighted by effect size that are commonly used to predict disease risk. For complex heritable diseases such as late onset Alzheimer’s disease (LOAD), PRS models fail to capture much of the heritability. Additionally, PRS models are highly dependent on the population structure of data on which effect sizes are assessed, and have poor generalizability to new data. OBJECTIVE. The goal of this study is to construct a paragenic risk score that, in addition to single genetic marker data used in PRS, incorporates epistatic interaction features and machine learning methods to predict lifetime risk for LOAD. METHODS. We construct a new state-of-the-art genetic model for lifetime risk of Alzheimer’s disease. Our approach innovates over PRS models in two ways: First, by directly incorporating epistatic interactions between SNP loci using an evolutionary algorithm guided by shared pathway information; and second, by estimating risk via an ensemble of machine learning models (gradient boosting machines and deep learning) instead of simple logistic regression. We compare the paragenic model to a PRS model from the literature trained on the same dataset. RESULTS. The paragenic model is significantly more accurate than the PRS model under 10-fold cross-validation, obtaining an AUC of 83% and near-clinically significant matched sensitivity/specificity of 75%, and remains significantly more accurate when evaluated on an independent holdout dataset. Additionally, the paragenic model maintains accuracy within APOE genotypes. CONCLUSION. Paragenic models show potential for improving lifetime disease risk prediction for complex heritable diseases such as LOAD over PRS models. Cold Spring Harbor Laboratory 2023-03-15 /pmc/articles/PMC9934790/ /pubmed/36798198 http://dx.doi.org/10.1101/2023.02.10.23285766 Text en https://creativecommons.org/licenses/by-nd/4.0/This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License (https://creativecommons.org/licenses/by-nd/4.0/) , which allows reusers to copy and distribute the material in any medium or format in unadapted form only, and only so long as attribution is given to the creator. The license allows for commercial use.
spellingShingle Article
Hermes, Stephen
Cady, Janet
Armentrout, Steven
O’Connor, James
Carlson, Sarah
Cruchaga, Carlos
Wingo, Thomas
Greytak, Ellen McRae
Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores
title Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores
title_full Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores
title_fullStr Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores
title_full_unstemmed Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores
title_short Epistatic Features and Machine Learning Improve Alzheimer’s Risk Prediction Over Polygenic Risk Scores
title_sort epistatic features and machine learning improve alzheimer’s risk prediction over polygenic risk scores
topic Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9934790/
https://www.ncbi.nlm.nih.gov/pubmed/36798198
http://dx.doi.org/10.1101/2023.02.10.23285766
work_keys_str_mv AT hermesstephen epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT cadyjanet epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT armentroutsteven epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT oconnorjames epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT carlsonsarah epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT cruchagacarlos epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT wingothomas epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT greytakellenmcrae epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores
AT epistaticfeaturesandmachinelearningimprovealzheimersriskpredictionoverpolygenicriskscores