Cargando…

Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks

Post-translational modifications (PTMs) provide an extensible framework for regulation of protein behavior beyond the diversity represented within the genome alone. While the rate of identification of PTMs has rapidly increased in recent years, our knowledge of PTM functionality encompasses less tha...

Descripción completa

Detalles Bibliográficos
Autores principales: Dewhurst, Henry M., Torres, Matthew P.
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Public Library of Science 2017
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5321281/
https://www.ncbi.nlm.nih.gov/pubmed/28225828
http://dx.doi.org/10.1371/journal.pone.0172572
_version_ 1782509665678000128
author Dewhurst, Henry M.
Torres, Matthew P.
author_facet Dewhurst, Henry M.
Torres, Matthew P.
author_sort Dewhurst, Henry M.
collection PubMed
description Post-translational modifications (PTMs) provide an extensible framework for regulation of protein behavior beyond the diversity represented within the genome alone. While the rate of identification of PTMs has rapidly increased in recent years, our knowledge of PTM functionality encompasses less than 5% of this data. We previously developed SAPH-ire (Structural Analysis of PTM Hotspots) for the prioritization of eukaryotic PTMs based on function potential of discrete modified alignment positions (MAPs) in a set of 8 protein families. A proteome-wide expansion of the dataset to all families of PTM-bearing, eukaryotic proteins with a representational crystal structure and the application of artificial neural network (ANN) models demonstrated the broader applicability of this approach. Although structural features of proteins have been repeatedly demonstrated to be predictive of PTM functionality, the availability of adequately resolved 3D structures in the Protein Data Bank (PDB) limits the scope of these methods. In order to bridge this gap and capture the larger set of PTM-bearing proteins without an available, homologous structure, we explored all available MAP features as ANN inputs to identify predictive models that do not rely on 3D protein structural data. This systematic, algorithmic approach explores 8 available input features in exhaustive combinations (247 models; size 2–8). To control for potential bias in random sampling for holdback in training sets, we iterated each model across 100 randomized, sample training and testing sets—yielding 24,700 individual ANNs. The size of the analyzed dataset and iterative generation of ANNs represents the largest and most thorough investigation of predictive models for PTM functionality to date. Comparison of input layer combinations allows us to quantify ANN performance with a high degree of confidence and subsequently select a top-ranked, robust fit model which highlights 3,687 MAPs, including 10,933 PTMs with a high probability of biological impact but without a currently known functional role.
format Online
Article
Text
id pubmed-5321281
institution National Center for Biotechnology Information
language English
publishDate 2017
publisher Public Library of Science
record_format MEDLINE/PubMed
spelling pubmed-53212812017-03-09 Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks Dewhurst, Henry M. Torres, Matthew P. PLoS One Research Article Post-translational modifications (PTMs) provide an extensible framework for regulation of protein behavior beyond the diversity represented within the genome alone. While the rate of identification of PTMs has rapidly increased in recent years, our knowledge of PTM functionality encompasses less than 5% of this data. We previously developed SAPH-ire (Structural Analysis of PTM Hotspots) for the prioritization of eukaryotic PTMs based on function potential of discrete modified alignment positions (MAPs) in a set of 8 protein families. A proteome-wide expansion of the dataset to all families of PTM-bearing, eukaryotic proteins with a representational crystal structure and the application of artificial neural network (ANN) models demonstrated the broader applicability of this approach. Although structural features of proteins have been repeatedly demonstrated to be predictive of PTM functionality, the availability of adequately resolved 3D structures in the Protein Data Bank (PDB) limits the scope of these methods. In order to bridge this gap and capture the larger set of PTM-bearing proteins without an available, homologous structure, we explored all available MAP features as ANN inputs to identify predictive models that do not rely on 3D protein structural data. This systematic, algorithmic approach explores 8 available input features in exhaustive combinations (247 models; size 2–8). To control for potential bias in random sampling for holdback in training sets, we iterated each model across 100 randomized, sample training and testing sets—yielding 24,700 individual ANNs. The size of the analyzed dataset and iterative generation of ANNs represents the largest and most thorough investigation of predictive models for PTM functionality to date. Comparison of input layer combinations allows us to quantify ANN performance with a high degree of confidence and subsequently select a top-ranked, robust fit model which highlights 3,687 MAPs, including 10,933 PTMs with a high probability of biological impact but without a currently known functional role. Public Library of Science 2017-02-22 /pmc/articles/PMC5321281/ /pubmed/28225828 http://dx.doi.org/10.1371/journal.pone.0172572 Text en © 2017 Dewhurst, Torres http://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/) , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
spellingShingle Research Article
Dewhurst, Henry M.
Torres, Matthew P.
Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks
title Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks
title_full Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks
title_fullStr Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks
title_full_unstemmed Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks
title_short Systematic analysis of non-structural protein features for the prediction of PTM function potential by artificial neural networks
title_sort systematic analysis of non-structural protein features for the prediction of ptm function potential by artificial neural networks
topic Research Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5321281/
https://www.ncbi.nlm.nih.gov/pubmed/28225828
http://dx.doi.org/10.1371/journal.pone.0172572
work_keys_str_mv AT dewhursthenrym systematicanalysisofnonstructuralproteinfeaturesforthepredictionofptmfunctionpotentialbyartificialneuralnetworks
AT torresmatthewp systematicanalysisofnonstructuralproteinfeaturesforthepredictionofptmfunctionpotentialbyartificialneuralnetworks