Cargando…
Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference
Accurate phylogenies are critical to taxonomy as well as studies of speciation processes and other evolutionary patterns. Accurate branch lengths in phylogenies are critical for dating and rate measurements. Such accuracy may be jeopardized by unacknowledged sequencing error. We use simulated data t...
Autores principales: | , |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
Genetics Society of America
2014
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4267948/ https://www.ncbi.nlm.nih.gov/pubmed/25378476 http://dx.doi.org/10.1534/g3.114.014365 |
_version_ | 1782349217866448896 |
---|---|
author | Kuhner, Mary K. McGill, James |
author_facet | Kuhner, Mary K. McGill, James |
author_sort | Kuhner, Mary K. |
collection | PubMed |
description | Accurate phylogenies are critical to taxonomy as well as studies of speciation processes and other evolutionary patterns. Accurate branch lengths in phylogenies are critical for dating and rate measurements. Such accuracy may be jeopardized by unacknowledged sequencing error. We use simulated data to test a correction for DNA sequencing error in maximum likelihood phylogeny inference. Over a wide range of data polymorphism and true error rate, we found that correcting for sequencing error improves recovery of the branch lengths, even if the assumed error rate is up to twice the true error rate. Low error rates have little effect on recovery of the topology. When error is high, correction improves topological inference; however, when error is extremely high, using an assumed error rate greater than the true error rate leads to poor recovery of both topology and branch lengths. The error correction approach tested here was proposed in 2004 but has not been widely used, perhaps because researchers do not want to commit to an estimate of the error rate. This study shows that correction with an approximate error rate is generally preferable to ignoring the issue. |
format | Online Article Text |
id | pubmed-4267948 |
institution | National Center for Biotechnology Information |
language | English |
publishDate | 2014 |
publisher | Genetics Society of America |
record_format | MEDLINE/PubMed |
spelling | pubmed-42679482014-12-23 Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference Kuhner, Mary K. McGill, James G3 (Bethesda) Investigations Accurate phylogenies are critical to taxonomy as well as studies of speciation processes and other evolutionary patterns. Accurate branch lengths in phylogenies are critical for dating and rate measurements. Such accuracy may be jeopardized by unacknowledged sequencing error. We use simulated data to test a correction for DNA sequencing error in maximum likelihood phylogeny inference. Over a wide range of data polymorphism and true error rate, we found that correcting for sequencing error improves recovery of the branch lengths, even if the assumed error rate is up to twice the true error rate. Low error rates have little effect on recovery of the topology. When error is high, correction improves topological inference; however, when error is extremely high, using an assumed error rate greater than the true error rate leads to poor recovery of both topology and branch lengths. The error correction approach tested here was proposed in 2004 but has not been widely used, perhaps because researchers do not want to commit to an estimate of the error rate. This study shows that correction with an approximate error rate is generally preferable to ignoring the issue. Genetics Society of America 2014-11-04 /pmc/articles/PMC4267948/ /pubmed/25378476 http://dx.doi.org/10.1534/g3.114.014365 Text en Copyright © 2014 Kuhner and McGill http://creativecommons.org/licenses/by/3.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution Unported License (http://creativecommons.org/licenses/by/3.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. |
spellingShingle | Investigations Kuhner, Mary K. McGill, James Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference |
title | Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference |
title_full | Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference |
title_fullStr | Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference |
title_full_unstemmed | Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference |
title_short | Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference |
title_sort | correcting for sequencing error in maximum likelihood phylogeny inference |
topic | Investigations |
url | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4267948/ https://www.ncbi.nlm.nih.gov/pubmed/25378476 http://dx.doi.org/10.1534/g3.114.014365 |
work_keys_str_mv | AT kuhnermaryk correctingforsequencingerrorinmaximumlikelihoodphylogenyinference AT mcgilljames correctingforsequencingerrorinmaximumlikelihoodphylogenyinference |