Cargando…

An automated framework for evaluation of deep learning models for splice site predictions

A novel framework for the automated evaluation of various deep learning-based splice site detectors is presented. The framework eliminates time-consuming development and experimenting activities for different codebases, architectures, and configurations to obtain the best models for a given RNA spli...

Descripción completa

Detalles Bibliográficos
Autores principales:	Zabardast, Amin, Tamer, Elif Güney, Son, Yeşim Aydın, Yılmaz, Arif
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Nature Publishing Group UK 2023
Materias:	Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10290104/ https://www.ncbi.nlm.nih.gov/pubmed/37353532 http://dx.doi.org/10.1038/s41598-023-34795-4

_version_	1785062420097531904
author	Zabardast, Amin Tamer, Elif Güney Son, Yeşim Aydın Yılmaz, Arif
author_facet	Zabardast, Amin Tamer, Elif Güney Son, Yeşim Aydın Yılmaz, Arif
author_sort	Zabardast, Amin
collection	PubMed
description	A novel framework for the automated evaluation of various deep learning-based splice site detectors is presented. The framework eliminates time-consuming development and experimenting activities for different codebases, architectures, and configurations to obtain the best models for a given RNA splice site dataset. RNA splicing is a cellular process in which pre-mRNAs are processed into mature mRNAs and used to produce multiple mRNA transcripts from a single gene sequence. Since the advancement of sequencing technologies, many splice site variants have been identified and associated with the diseases. So, RNA splice site prediction is essential for gene finding, genome annotation, disease-causing variants, and identification of potential biomarkers. Recently, deep learning models performed highly accurately for classifying genomic signals. Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) and its bidirectional version (BLSTM), Gated Recurrent Unit (GRU), and its bidirectional version (BGRU) are promising models. During genomic data analysis, CNN’s locality feature helps where each nucleotide correlates with other bases in its vicinity. In contrast, BLSTM can be trained bidirectionally, allowing sequential data to be processed from forward and reverse directions. Therefore, it can process 1-D encoded genomic data effectively. Even though both methods have been used in the literature, a performance comparison was missing. To compare selected models under similar conditions, we have created a blueprint for a series of networks with five different levels. As a case study, we compared CNN and BLSTM models’ learning capabilities as building blocks for RNA splice site prediction in two different datasets. Overall, CNN performed better with [Formula: see text] accuracy ([Formula: see text] improvement), [Formula: see text] F1 score ([Formula: see text] improvement), and [Formula: see text] AUC-PR ([Formula: see text] improvement) in human splice site prediction. Likewise, an outperforming performance with [Formula: see text] accuracy ([Formula: see text] improvement), [Formula: see text] F1 score ([Formula: see text] improvement), and [Formula: see text] AUC-PR ([Formula: see text] improvement) is achieved in C. elegans splice site prediction. Overall, our results showed that CNN learns faster than BLSTM and BGRU. Moreover, CNN performs better at extracting sequence patterns than BLSTM and BGRU. To our knowledge, no other framework is developed explicitly for evaluating splice detection models to decide the best possible model in an automated manner. So, the proposed framework and the blueprint would help selecting different deep learning models, such as CNN vs. BLSTM and BGRU, for splice site analysis or similar classification tasks and in different problems.
format	Online Article Text
id	pubmed-10290104
institution	National Center for Biotechnology Information
language	English
publishDate	2023
publisher	Nature Publishing Group UK
record_format	MEDLINE/PubMed
spelling	pubmed-102901042023-06-25 An automated framework for evaluation of deep learning models for splice site predictions Zabardast, Amin Tamer, Elif Güney Son, Yeşim Aydın Yılmaz, Arif Sci Rep Article A novel framework for the automated evaluation of various deep learning-based splice site detectors is presented. The framework eliminates time-consuming development and experimenting activities for different codebases, architectures, and configurations to obtain the best models for a given RNA splice site dataset. RNA splicing is a cellular process in which pre-mRNAs are processed into mature mRNAs and used to produce multiple mRNA transcripts from a single gene sequence. Since the advancement of sequencing technologies, many splice site variants have been identified and associated with the diseases. So, RNA splice site prediction is essential for gene finding, genome annotation, disease-causing variants, and identification of potential biomarkers. Recently, deep learning models performed highly accurately for classifying genomic signals. Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) and its bidirectional version (BLSTM), Gated Recurrent Unit (GRU), and its bidirectional version (BGRU) are promising models. During genomic data analysis, CNN’s locality feature helps where each nucleotide correlates with other bases in its vicinity. In contrast, BLSTM can be trained bidirectionally, allowing sequential data to be processed from forward and reverse directions. Therefore, it can process 1-D encoded genomic data effectively. Even though both methods have been used in the literature, a performance comparison was missing. To compare selected models under similar conditions, we have created a blueprint for a series of networks with five different levels. As a case study, we compared CNN and BLSTM models’ learning capabilities as building blocks for RNA splice site prediction in two different datasets. Overall, CNN performed better with [Formula: see text] accuracy ([Formula: see text] improvement), [Formula: see text] F1 score ([Formula: see text] improvement), and [Formula: see text] AUC-PR ([Formula: see text] improvement) in human splice site prediction. Likewise, an outperforming performance with [Formula: see text] accuracy ([Formula: see text] improvement), [Formula: see text] F1 score ([Formula: see text] improvement), and [Formula: see text] AUC-PR ([Formula: see text] improvement) is achieved in C. elegans splice site prediction. Overall, our results showed that CNN learns faster than BLSTM and BGRU. Moreover, CNN performs better at extracting sequence patterns than BLSTM and BGRU. To our knowledge, no other framework is developed explicitly for evaluating splice detection models to decide the best possible model in an automated manner. So, the proposed framework and the blueprint would help selecting different deep learning models, such as CNN vs. BLSTM and BGRU, for splice site analysis or similar classification tasks and in different problems. Nature Publishing Group UK 2023-06-23 /pmc/articles/PMC10290104/ /pubmed/37353532 http://dx.doi.org/10.1038/s41598-023-34795-4 Text en © The Author(s) 2023 https://creativecommons.org/licenses/by/4.0/Open AccessThis article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/ (https://creativecommons.org/licenses/by/4.0/) .
spellingShingle	Article Zabardast, Amin Tamer, Elif Güney Son, Yeşim Aydın Yılmaz, Arif An automated framework for evaluation of deep learning models for splice site predictions
title	An automated framework for evaluation of deep learning models for splice site predictions
title_full	An automated framework for evaluation of deep learning models for splice site predictions
title_fullStr	An automated framework for evaluation of deep learning models for splice site predictions
title_full_unstemmed	An automated framework for evaluation of deep learning models for splice site predictions
title_short	An automated framework for evaluation of deep learning models for splice site predictions
title_sort	automated framework for evaluation of deep learning models for splice site predictions
topic	Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10290104/ https://www.ncbi.nlm.nih.gov/pubmed/37353532 http://dx.doi.org/10.1038/s41598-023-34795-4
work_keys_str_mv	AT zabardastamin anautomatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT tamerelifguney anautomatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT sonyesimaydın anautomatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT yılmazarif anautomatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT zabardastamin automatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT tamerelifguney automatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT sonyesimaydın automatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions AT yılmazarif automatedframeworkforevaluationofdeeplearningmodelsforsplicesitepredictions

An automated framework for evaluation of deep learning models for splice site predictions

Ejemplares similares