Cargando…

Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning

Proteomics instrumentation and the corresponding bioinformatics tools have evolved at a rapid pace in the last 20 years, whereas the exploitation of deep learning techniques in proteomics is on the horizon. The ability to revisit proteomics raw data, in particular, could be a valuable resource for m...

Descripción completa

Detalles Bibliográficos
Autores principales:	Rafay, Abdul, Aziz, Muzzamil, Zia, Amjad, Asif, Abdul R.
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	MDPI 2023
Materias:	Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10222177/ https://www.ncbi.nlm.nih.gov/pubmed/37240960 http://dx.doi.org/10.3390/jpm13050790

_version_	1785049634884812800
author	Rafay, Abdul Aziz, Muzzamil Zia, Amjad Asif, Abdul R.
author_facet	Rafay, Abdul Aziz, Muzzamil Zia, Amjad Asif, Abdul R.
author_sort	Rafay, Abdul
collection	PubMed
description	Proteomics instrumentation and the corresponding bioinformatics tools have evolved at a rapid pace in the last 20 years, whereas the exploitation of deep learning techniques in proteomics is on the horizon. The ability to revisit proteomics raw data, in particular, could be a valuable resource for machine learning applications seeking new insight into protein expression and functions of previously acquired data from different instruments under various lab conditions. We map publicly available proteomics repositories (such as ProteomeXchange) and relevant publications to extract MS/MS data to form one large database that contains the patient history and mass spectrometric data acquired for the patient sample. The extracted mapped dataset should enable the research to overcome the issues attached to the dispersions of proteomics data on the internet, which makes it difficult to apply emerging new bioinformatics tools and deep learning algorithms. The workflow proposed in this study enables a linked large dataset of heart-related proteomics data, which could be easily and efficiently applied to machine learning and deep learning algorithms for futuristic predictions of heart diseases and modeling. Data scraping and crawling offer a powerful tool to harvest and prepare the training and test datasets; however, the authors advocate caution because of ethical and legal issues, as well as the need to ensure the quality and accuracy of the data that are being collected.
format	Online Article Text
id	pubmed-10222177
institution	National Center for Biotechnology Information
language	English
publishDate	2023
publisher	MDPI
record_format	MEDLINE/PubMed
spelling	pubmed-102221772023-05-28 Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning Rafay, Abdul Aziz, Muzzamil Zia, Amjad Asif, Abdul R. J Pers Med Article Proteomics instrumentation and the corresponding bioinformatics tools have evolved at a rapid pace in the last 20 years, whereas the exploitation of deep learning techniques in proteomics is on the horizon. The ability to revisit proteomics raw data, in particular, could be a valuable resource for machine learning applications seeking new insight into protein expression and functions of previously acquired data from different instruments under various lab conditions. We map publicly available proteomics repositories (such as ProteomeXchange) and relevant publications to extract MS/MS data to form one large database that contains the patient history and mass spectrometric data acquired for the patient sample. The extracted mapped dataset should enable the research to overcome the issues attached to the dispersions of proteomics data on the internet, which makes it difficult to apply emerging new bioinformatics tools and deep learning algorithms. The workflow proposed in this study enables a linked large dataset of heart-related proteomics data, which could be easily and efficiently applied to machine learning and deep learning algorithms for futuristic predictions of heart diseases and modeling. Data scraping and crawling offer a powerful tool to harvest and prepare the training and test datasets; however, the authors advocate caution because of ethical and legal issues, as well as the need to ensure the quality and accuracy of the data that are being collected. MDPI 2023-05-02 /pmc/articles/PMC10222177/ /pubmed/37240960 http://dx.doi.org/10.3390/jpm13050790 Text en © 2023 by the authors. https://creativecommons.org/licenses/by/4.0/Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
spellingShingle	Article Rafay, Abdul Aziz, Muzzamil Zia, Amjad Asif, Abdul R. Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning
title	Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning
title_full	Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning
title_fullStr	Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning
title_full_unstemmed	Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning
title_short	Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning
title_sort	automated retrieval of heterogeneous proteomic data for machine learning
topic	Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10222177/ https://www.ncbi.nlm.nih.gov/pubmed/37240960 http://dx.doi.org/10.3390/jpm13050790
work_keys_str_mv	AT rafayabdul automatedretrievalofheterogeneousproteomicdataformachinelearning AT azizmuzzamil automatedretrievalofheterogeneousproteomicdataformachinelearning AT ziaamjad automatedretrievalofheterogeneousproteomicdataformachinelearning AT asifabdulr automatedretrievalofheterogeneousproteomicdataformachinelearning

Automated Retrieval of Heterogeneous Proteomic Data for Machine Learning

Ejemplares similares