Cargando…

Chromatin accessibility prediction via a hybrid deep convolutional neural network

MOTIVATION: A majority of known genetic variants associated with human-inherited diseases lie in non-coding regions that lack adequate interpretation, making it indispensable to systematically discover functional sites at the whole genome level and precisely decipher their implications in a comprehe...

Descripción completa

Detalles Bibliográficos
Autores principales: Liu, Qiao, Xia, Fei, Yin, Qijin, Jiang, Rui
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Oxford University Press 2018
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6192215/
https://www.ncbi.nlm.nih.gov/pubmed/29069282
http://dx.doi.org/10.1093/bioinformatics/btx679
_version_ 1783363868466085888
author Liu, Qiao
Xia, Fei
Yin, Qijin
Jiang, Rui
author_facet Liu, Qiao
Xia, Fei
Yin, Qijin
Jiang, Rui
author_sort Liu, Qiao
collection PubMed
description MOTIVATION: A majority of known genetic variants associated with human-inherited diseases lie in non-coding regions that lack adequate interpretation, making it indispensable to systematically discover functional sites at the whole genome level and precisely decipher their implications in a comprehensive manner. Although computational approaches have been complementing high-throughput biological experiments towards the annotation of the human genome, it still remains a big challenge to accurately annotate regulatory elements in the context of a specific cell type via automatic learning of the DNA sequence code from large-scale sequencing data. Indeed, the development of an accurate and interpretable model to learn the DNA sequence signature and further enable the identification of causative genetic variants has become essential in both genomic and genetic studies. RESULTS: We proposed Deopen, a hybrid framework mainly based on a deep convolutional neural network, to automatically learn the regulatory code of DNA sequences and predict chromatin accessibility. In a series of comparison with existing methods, we show the superior performance of our model in not only the classification of accessible regions against background sequences sampled at random, but also the regression of DNase-seq signals. Besides, we further visualize the convolutional kernels and show the match of identified sequence signatures and known motifs. We finally demonstrate the sensitivity of our model in finding causative noncoding variants in the analysis of a breast cancer dataset. We expect to see wide applications of Deopen with either public or in-house chromatin accessibility data in the annotation of the human genome and the identification of non-coding variants associated with diseases. AVAILABILITY AND IMPLEMENTATION: Deopen is freely available at https://github.com/kimmo1019/Deopen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
format Online
Article
Text
id pubmed-6192215
institution National Center for Biotechnology Information
language English
publishDate 2018
publisher Oxford University Press
record_format MEDLINE/PubMed
spelling pubmed-61922152019-03-01 Chromatin accessibility prediction via a hybrid deep convolutional neural network Liu, Qiao Xia, Fei Yin, Qijin Jiang, Rui Bioinformatics Original Papers MOTIVATION: A majority of known genetic variants associated with human-inherited diseases lie in non-coding regions that lack adequate interpretation, making it indispensable to systematically discover functional sites at the whole genome level and precisely decipher their implications in a comprehensive manner. Although computational approaches have been complementing high-throughput biological experiments towards the annotation of the human genome, it still remains a big challenge to accurately annotate regulatory elements in the context of a specific cell type via automatic learning of the DNA sequence code from large-scale sequencing data. Indeed, the development of an accurate and interpretable model to learn the DNA sequence signature and further enable the identification of causative genetic variants has become essential in both genomic and genetic studies. RESULTS: We proposed Deopen, a hybrid framework mainly based on a deep convolutional neural network, to automatically learn the regulatory code of DNA sequences and predict chromatin accessibility. In a series of comparison with existing methods, we show the superior performance of our model in not only the classification of accessible regions against background sequences sampled at random, but also the regression of DNase-seq signals. Besides, we further visualize the convolutional kernels and show the match of identified sequence signatures and known motifs. We finally demonstrate the sensitivity of our model in finding causative noncoding variants in the analysis of a breast cancer dataset. We expect to see wide applications of Deopen with either public or in-house chromatin accessibility data in the annotation of the human genome and the identification of non-coding variants associated with diseases. AVAILABILITY AND IMPLEMENTATION: Deopen is freely available at https://github.com/kimmo1019/Deopen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Oxford University Press 2018-03-01 2017-10-23 /pmc/articles/PMC6192215/ /pubmed/29069282 http://dx.doi.org/10.1093/bioinformatics/btx679 Text en © The Author 2017. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com https://creativecommons.org/licenses/by/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle Original Papers
Liu, Qiao
Xia, Fei
Yin, Qijin
Jiang, Rui
Chromatin accessibility prediction via a hybrid deep convolutional neural network
title Chromatin accessibility prediction via a hybrid deep convolutional neural network
title_full Chromatin accessibility prediction via a hybrid deep convolutional neural network
title_fullStr Chromatin accessibility prediction via a hybrid deep convolutional neural network
title_full_unstemmed Chromatin accessibility prediction via a hybrid deep convolutional neural network
title_short Chromatin accessibility prediction via a hybrid deep convolutional neural network
title_sort chromatin accessibility prediction via a hybrid deep convolutional neural network
topic Original Papers
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6192215/
https://www.ncbi.nlm.nih.gov/pubmed/29069282
http://dx.doi.org/10.1093/bioinformatics/btx679
work_keys_str_mv AT liuqiao chromatinaccessibilitypredictionviaahybriddeepconvolutionalneuralnetwork
AT xiafei chromatinaccessibilitypredictionviaahybriddeepconvolutionalneuralnetwork
AT yinqijin chromatinaccessibilitypredictionviaahybriddeepconvolutionalneuralnetwork
AT jiangrui chromatinaccessibilitypredictionviaahybriddeepconvolutionalneuralnetwork