Cargando…

Finding Protein-Coding Genes through Human Polymorphisms

Human gene catalogs are fundamental to the study of human biology and medicine. But they are all based on open reading frames (ORFs) in a reference genome sequence (with allowance for introns). Individual genomes, however, are polymorphic: their sequences are not identical. There has been much resea...

Descripción completa

Detalles Bibliográficos
Autores principales: Wijaya, Edward, Frith, Martin C., Horton, Paul, Asai, Kiyoshi
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Public Library of Science 2013
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3551959/
https://www.ncbi.nlm.nih.gov/pubmed/23349826
http://dx.doi.org/10.1371/journal.pone.0054210
_version_ 1782256651661737984
author Wijaya, Edward
Frith, Martin C.
Horton, Paul
Asai, Kiyoshi
author_facet Wijaya, Edward
Frith, Martin C.
Horton, Paul
Asai, Kiyoshi
author_sort Wijaya, Edward
collection PubMed
description Human gene catalogs are fundamental to the study of human biology and medicine. But they are all based on open reading frames (ORFs) in a reference genome sequence (with allowance for introns). Individual genomes, however, are polymorphic: their sequences are not identical. There has been much research on how polymorphism affects previously-identified genes, but no research has been done on how it affects gene identification itself. We computationally predict protein-coding genes in a straightforward manner, by finding long ORFs in mRNA sequences aligned to the reference genome. We systematically test the effect of known polymorphisms with this procedure. Polymorphisms can not only disrupt ORFs, they can also create long ORFs that do not exist in the reference sequence. We found 5,737 putative protein-coding genes that do not exist in the reference, whose protein-coding status is supported by homology to known proteins. On average 10% of these genes are located in the genomic regions devoid of annotated genes in 12 other catalogs. Our statistical analysis showed that these ORFs are unlikely to occur by chance.
format Online
Article
Text
id pubmed-3551959
institution National Center for Biotechnology Information
language English
publishDate 2013
publisher Public Library of Science
record_format MEDLINE/PubMed
spelling pubmed-35519592013-01-24 Finding Protein-Coding Genes through Human Polymorphisms Wijaya, Edward Frith, Martin C. Horton, Paul Asai, Kiyoshi PLoS One Research Article Human gene catalogs are fundamental to the study of human biology and medicine. But they are all based on open reading frames (ORFs) in a reference genome sequence (with allowance for introns). Individual genomes, however, are polymorphic: their sequences are not identical. There has been much research on how polymorphism affects previously-identified genes, but no research has been done on how it affects gene identification itself. We computationally predict protein-coding genes in a straightforward manner, by finding long ORFs in mRNA sequences aligned to the reference genome. We systematically test the effect of known polymorphisms with this procedure. Polymorphisms can not only disrupt ORFs, they can also create long ORFs that do not exist in the reference sequence. We found 5,737 putative protein-coding genes that do not exist in the reference, whose protein-coding status is supported by homology to known proteins. On average 10% of these genes are located in the genomic regions devoid of annotated genes in 12 other catalogs. Our statistical analysis showed that these ORFs are unlikely to occur by chance. Public Library of Science 2013-01-22 /pmc/articles/PMC3551959/ /pubmed/23349826 http://dx.doi.org/10.1371/journal.pone.0054210 Text en © 2013 Wijaya et al http://creativecommons.org/licenses/by/4.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are properly credited.
spellingShingle Research Article
Wijaya, Edward
Frith, Martin C.
Horton, Paul
Asai, Kiyoshi
Finding Protein-Coding Genes through Human Polymorphisms
title Finding Protein-Coding Genes through Human Polymorphisms
title_full Finding Protein-Coding Genes through Human Polymorphisms
title_fullStr Finding Protein-Coding Genes through Human Polymorphisms
title_full_unstemmed Finding Protein-Coding Genes through Human Polymorphisms
title_short Finding Protein-Coding Genes through Human Polymorphisms
title_sort finding protein-coding genes through human polymorphisms
topic Research Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3551959/
https://www.ncbi.nlm.nih.gov/pubmed/23349826
http://dx.doi.org/10.1371/journal.pone.0054210
work_keys_str_mv AT wijayaedward findingproteincodinggenesthroughhumanpolymorphisms
AT frithmartinc findingproteincodinggenesthroughhumanpolymorphisms
AT hortonpaul findingproteincodinggenesthroughhumanpolymorphisms
AT asaikiyoshi findingproteincodinggenesthroughhumanpolymorphisms