Cargando…

Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes

We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference-based data formats do not fully exploit the redundancy in po...

Descripción completa

Detalles Bibliográficos
Autores principales:	Dolle, Dirk D., Liu, Zhicheng, Cotten, Matthew, Simpson, Jared T., Iqbal, Zamin, Durbin, Richard, McCarthy, Shane A., Keane, Thomas M.
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Cold Spring Harbor Laboratory Press 2017
Materias:	Method
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5287235/ https://www.ncbi.nlm.nih.gov/pubmed/27986821 http://dx.doi.org/10.1101/gr.211748.116

_version_	1782504132583620608
author	Dolle, Dirk D. Liu, Zhicheng Cotten, Matthew Simpson, Jared T. Iqbal, Zamin Durbin, Richard McCarthy, Shane A. Keane, Thomas M.
author_facet	Dolle, Dirk D. Liu, Zhicheng Cotten, Matthew Simpson, Jared T. Iqbal, Zamin Durbin, Richard McCarthy, Shane A. Keane, Thomas M.
author_sort	Dolle, Dirk D.
collection	PubMed
description	We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference-based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows–Wheeler transform (BWT) and FM-index have been widely employed as a full-text searchable index for read alignment and de novo assembly. We introduce the concept of a population BWT and use it to store and index the sequencing reads of 2705 samples from the 1000 Genomes Project. A key feature is that, as more genomes are added, identical read sequences are increasingly observed, and compression becomes more efficient. We assess the support in the 1000 Genomes read data for every base position of two human reference assembly versions, identifying that 3.2 Mbp with population support was lost in the transition from GRCh37 with 13.7 Mbp added to GRCh38. We show that the vast majority of variant alleles can be uniquely described by overlapping 31-mers and show how rapid and accurate SNP and indel genotyping can be carried out across the genomes in the population BWT. We use the population BWT to carry out nonreference queries to search for the presence of all known viral genomes and discover human T-lymphotropic virus 1 integrations in six samples in a recognized epidemiological distribution.
format	Online Article Text
id	pubmed-5287235
institution	National Center for Biotechnology Information
language	English
publishDate	2017
publisher	Cold Spring Harbor Laboratory Press
record_format	MEDLINE/PubMed
spelling	pubmed-52872352017-08-01 Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes Dolle, Dirk D. Liu, Zhicheng Cotten, Matthew Simpson, Jared T. Iqbal, Zamin Durbin, Richard McCarthy, Shane A. Keane, Thomas M. Genome Res Method We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference-based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows–Wheeler transform (BWT) and FM-index have been widely employed as a full-text searchable index for read alignment and de novo assembly. We introduce the concept of a population BWT and use it to store and index the sequencing reads of 2705 samples from the 1000 Genomes Project. A key feature is that, as more genomes are added, identical read sequences are increasingly observed, and compression becomes more efficient. We assess the support in the 1000 Genomes read data for every base position of two human reference assembly versions, identifying that 3.2 Mbp with population support was lost in the transition from GRCh37 with 13.7 Mbp added to GRCh38. We show that the vast majority of variant alleles can be uniquely described by overlapping 31-mers and show how rapid and accurate SNP and indel genotyping can be carried out across the genomes in the population BWT. We use the population BWT to carry out nonreference queries to search for the presence of all known viral genomes and discover human T-lymphotropic virus 1 integrations in six samples in a recognized epidemiological distribution. Cold Spring Harbor Laboratory Press 2017-02 /pmc/articles/PMC5287235/ /pubmed/27986821 http://dx.doi.org/10.1101/gr.211748.116 Text en © 2017 Dolle et al.; Published by Cold Spring Harbor Laboratory Press http://creativecommons.org/licenses/by-nc/4.0/ This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first six months after the full-issue publication date (see http://genome.cshlp.org/site/misc/terms.xhtml). After six months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.
spellingShingle	Method Dolle, Dirk D. Liu, Zhicheng Cotten, Matthew Simpson, Jared T. Iqbal, Zamin Durbin, Richard McCarthy, Shane A. Keane, Thomas M. Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
title	Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
title_full	Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
title_fullStr	Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
title_full_unstemmed	Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
title_short	Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
title_sort	using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes
topic	Method
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5287235/ https://www.ncbi.nlm.nih.gov/pubmed/27986821 http://dx.doi.org/10.1101/gr.211748.116
work_keys_str_mv	AT dolledirkd usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT liuzhicheng usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT cottenmatthew usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT simpsonjaredt usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT iqbalzamin usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT durbinrichard usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT mccarthyshanea usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes AT keanethomasm usingreferencefreecompresseddatastructurestoanalyzesequencingreadsfromthousandsofhumangenomes

Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes

Ejemplares similares