Cargando…

nDNA-prot: identification of DNA-binding proteins based on unbalanced classification

BACKGROUND: DNA-binding proteins are vital for the study of cellular processes. In recent genome engineering studies, the identification of proteins with certain functions has become increasingly important and needs to be performed rapidly and efficiently. In previous years, several approaches have...

Descripción completa

Detalles Bibliográficos
Autores principales:	Song, Li, Li, Dapeng, Zeng, Xiangxiang, Wu, Yunfeng, Guo, Li, Zou, Quan
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	BioMed Central 2014
Materias:	Research Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4165999/ https://www.ncbi.nlm.nih.gov/pubmed/25196432 http://dx.doi.org/10.1186/1471-2105-15-298

_version_	1782335182779449344
author	Song, Li Li, Dapeng Zeng, Xiangxiang Wu, Yunfeng Guo, Li Zou, Quan
author_facet	Song, Li Li, Dapeng Zeng, Xiangxiang Wu, Yunfeng Guo, Li Zou, Quan
author_sort	Song, Li
collection	PubMed
description	BACKGROUND: DNA-binding proteins are vital for the study of cellular processes. In recent genome engineering studies, the identification of proteins with certain functions has become increasingly important and needs to be performed rapidly and efficiently. In previous years, several approaches have been developed to improve the identification of DNA-binding proteins. However, the currently available resources are insufficient to accurately identify these proteins. Because of this, the previous research has been limited by the relatively unbalanced accuracy rate and the low identification success of the current methods. RESULTS: In this paper, we explored the practicality of modelling DNA binding identification and simultaneously employed an ensemble classifier, and a new predictor (nDNA-Prot) was designed. The presented framework is comprised of two stages: a 188-dimension feature extraction method to obtain the protein structure and an ensemble classifier designated as imDC. Experiments using different datasets showed that our method is more successful than the traditional methods in identifying DNA-binding proteins. The identification was conducted using a feature that selected the minimum Redundancy and Maximum Relevance (mRMR). An accuracy rate of 95.80% and an Area Under the Curve (AUC) value of 0.986 were obtained in a cross validation. A test dataset was tested in our method and resulted in an 86% accuracy, versus a 76% using iDNA-Prot and a 68% accuracy using DNA-Prot. CONCLUSIONS: Our method can help to accurately identify DNA-binding proteins, and the web server is accessible at http://datamining.xmu.edu.cn/~songli/nDNA. In addition, we also predicted possible DNA-binding protein sequences in all of the sequences from the UniProtKB/Swiss-Prot database. ELECTRONIC SUPPLEMENTARY MATERIAL: The online version of this article (doi:10.1186/1471-2105-15-298) contains supplementary material, which is available to authorized users.
format	Online Article Text
id	pubmed-4165999
institution	National Center for Biotechnology Information
language	English
publishDate	2014
publisher	BioMed Central
record_format	MEDLINE/PubMed
spelling	pubmed-41659992014-09-18 nDNA-prot: identification of DNA-binding proteins based on unbalanced classification Song, Li Li, Dapeng Zeng, Xiangxiang Wu, Yunfeng Guo, Li Zou, Quan BMC Bioinformatics Research Article BACKGROUND: DNA-binding proteins are vital for the study of cellular processes. In recent genome engineering studies, the identification of proteins with certain functions has become increasingly important and needs to be performed rapidly and efficiently. In previous years, several approaches have been developed to improve the identification of DNA-binding proteins. However, the currently available resources are insufficient to accurately identify these proteins. Because of this, the previous research has been limited by the relatively unbalanced accuracy rate and the low identification success of the current methods. RESULTS: In this paper, we explored the practicality of modelling DNA binding identification and simultaneously employed an ensemble classifier, and a new predictor (nDNA-Prot) was designed. The presented framework is comprised of two stages: a 188-dimension feature extraction method to obtain the protein structure and an ensemble classifier designated as imDC. Experiments using different datasets showed that our method is more successful than the traditional methods in identifying DNA-binding proteins. The identification was conducted using a feature that selected the minimum Redundancy and Maximum Relevance (mRMR). An accuracy rate of 95.80% and an Area Under the Curve (AUC) value of 0.986 were obtained in a cross validation. A test dataset was tested in our method and resulted in an 86% accuracy, versus a 76% using iDNA-Prot and a 68% accuracy using DNA-Prot. CONCLUSIONS: Our method can help to accurately identify DNA-binding proteins, and the web server is accessible at http://datamining.xmu.edu.cn/~songli/nDNA. In addition, we also predicted possible DNA-binding protein sequences in all of the sequences from the UniProtKB/Swiss-Prot database. ELECTRONIC SUPPLEMENTARY MATERIAL: The online version of this article (doi:10.1186/1471-2105-15-298) contains supplementary material, which is available to authorized users. BioMed Central 2014-09-08 /pmc/articles/PMC4165999/ /pubmed/25196432 http://dx.doi.org/10.1186/1471-2105-15-298 Text en © Song et al.; licensee BioMed Central Ltd. 2014 This article is published under license to BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly credited. The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in this article, unless otherwise stated.
spellingShingle	Research Article Song, Li Li, Dapeng Zeng, Xiangxiang Wu, Yunfeng Guo, Li Zou, Quan nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
title	nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
title_full	nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
title_fullStr	nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
title_full_unstemmed	nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
title_short	nDNA-prot: identification of DNA-binding proteins based on unbalanced classification
title_sort	ndna-prot: identification of dna-binding proteins based on unbalanced classification
topic	Research Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4165999/ https://www.ncbi.nlm.nih.gov/pubmed/25196432 http://dx.doi.org/10.1186/1471-2105-15-298
work_keys_str_mv	AT songli ndnaprotidentificationofdnabindingproteinsbasedonunbalancedclassification AT lidapeng ndnaprotidentificationofdnabindingproteinsbasedonunbalancedclassification AT zengxiangxiang ndnaprotidentificationofdnabindingproteinsbasedonunbalancedclassification AT wuyunfeng ndnaprotidentificationofdnabindingproteinsbasedonunbalancedclassification AT guoli ndnaprotidentificationofdnabindingproteinsbasedonunbalancedclassification AT zouquan ndnaprotidentificationofdnabindingproteinsbasedonunbalancedclassification

nDNA-prot: identification of DNA-binding proteins based on unbalanced classification

Ejemplares similares