Cargando…

A privacy-preserving solution for compressed storage and selective retrieval of genomic data

In clinical genomics, the continuous evolution of bioinformatic algorithms and sequencing platforms makes it beneficial to store patients’ complete aligned genomic data in addition to variant calls relative to a reference sequence. Due to the large size of human genome sequence data files (varying f...

Descripción completa

Detalles Bibliográficos
Autores principales: Huang, Zhicong, Ayday, Erman, Lin, Huang, Aiyar, Raeka S., Molyneaux, Adam, Xu, Zhenyu, Fellay, Jacques, Steinmetz, Lars M., Hubaux, Jean-Pierre
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Cold Spring Harbor Laboratory Press 2016
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5131820/
https://www.ncbi.nlm.nih.gov/pubmed/27789525
http://dx.doi.org/10.1101/gr.206870.116
_version_ 1782470951136395264
author Huang, Zhicong
Ayday, Erman
Lin, Huang
Aiyar, Raeka S.
Molyneaux, Adam
Xu, Zhenyu
Fellay, Jacques
Steinmetz, Lars M.
Hubaux, Jean-Pierre
author_facet Huang, Zhicong
Ayday, Erman
Lin, Huang
Aiyar, Raeka S.
Molyneaux, Adam
Xu, Zhenyu
Fellay, Jacques
Steinmetz, Lars M.
Hubaux, Jean-Pierre
author_sort Huang, Zhicong
collection PubMed
description In clinical genomics, the continuous evolution of bioinformatic algorithms and sequencing platforms makes it beneficial to store patients’ complete aligned genomic data in addition to variant calls relative to a reference sequence. Due to the large size of human genome sequence data files (varying from 30 GB to 200 GB depending on coverage), two major challenges facing genomics laboratories are the costs of storage and the efficiency of the initial data processing. In addition, privacy of genomic data is becoming an increasingly serious concern, yet no standard data storage solutions exist that enable compression, encryption, and selective retrieval. Here we present a privacy-preserving solution named SECRAM (Selective retrieval on Encrypted and Compressed Reference-oriented Alignment Map) for the secure storage of compressed aligned genomic data. Our solution enables selective retrieval of encrypted data and improves the efficiency of downstream analysis (e.g., variant calling). Compared with BAM, the de facto standard for storing aligned genomic data, SECRAM uses 18% less storage. Compared with CRAM, one of the most compressed nonencrypted formats (using 34% less storage than BAM), SECRAM maintains efficient compression and downstream data processing, while allowing for unprecedented levels of security in genomic data storage. Compared with previous work, the distinguishing features of SECRAM are that (1) it is position-based instead of read-based, and (2) it allows random querying of a subregion from a BAM-like file in an encrypted form. Our method thus offers a space-saving, privacy-preserving, and effective solution for the storage of clinical genomic data.
format Online
Article
Text
id pubmed-5131820
institution National Center for Biotechnology Information
language English
publishDate 2016
publisher Cold Spring Harbor Laboratory Press
record_format MEDLINE/PubMed
spelling pubmed-51318202017-06-01 A privacy-preserving solution for compressed storage and selective retrieval of genomic data Huang, Zhicong Ayday, Erman Lin, Huang Aiyar, Raeka S. Molyneaux, Adam Xu, Zhenyu Fellay, Jacques Steinmetz, Lars M. Hubaux, Jean-Pierre Genome Res Method In clinical genomics, the continuous evolution of bioinformatic algorithms and sequencing platforms makes it beneficial to store patients’ complete aligned genomic data in addition to variant calls relative to a reference sequence. Due to the large size of human genome sequence data files (varying from 30 GB to 200 GB depending on coverage), two major challenges facing genomics laboratories are the costs of storage and the efficiency of the initial data processing. In addition, privacy of genomic data is becoming an increasingly serious concern, yet no standard data storage solutions exist that enable compression, encryption, and selective retrieval. Here we present a privacy-preserving solution named SECRAM (Selective retrieval on Encrypted and Compressed Reference-oriented Alignment Map) for the secure storage of compressed aligned genomic data. Our solution enables selective retrieval of encrypted data and improves the efficiency of downstream analysis (e.g., variant calling). Compared with BAM, the de facto standard for storing aligned genomic data, SECRAM uses 18% less storage. Compared with CRAM, one of the most compressed nonencrypted formats (using 34% less storage than BAM), SECRAM maintains efficient compression and downstream data processing, while allowing for unprecedented levels of security in genomic data storage. Compared with previous work, the distinguishing features of SECRAM are that (1) it is position-based instead of read-based, and (2) it allows random querying of a subregion from a BAM-like file in an encrypted form. Our method thus offers a space-saving, privacy-preserving, and effective solution for the storage of clinical genomic data. Cold Spring Harbor Laboratory Press 2016-12 /pmc/articles/PMC5131820/ /pubmed/27789525 http://dx.doi.org/10.1101/gr.206870.116 Text en © 2016 Huang et al.; Published by Cold Spring Harbor Laboratory Press http://creativecommons.org/licenses/by-nc/4.0/ This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first six months after the full-issue publication date (see http://genome.cshlp.org/site/misc/terms.xhtml). After six months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.
spellingShingle Method
Huang, Zhicong
Ayday, Erman
Lin, Huang
Aiyar, Raeka S.
Molyneaux, Adam
Xu, Zhenyu
Fellay, Jacques
Steinmetz, Lars M.
Hubaux, Jean-Pierre
A privacy-preserving solution for compressed storage and selective retrieval of genomic data
title A privacy-preserving solution for compressed storage and selective retrieval of genomic data
title_full A privacy-preserving solution for compressed storage and selective retrieval of genomic data
title_fullStr A privacy-preserving solution for compressed storage and selective retrieval of genomic data
title_full_unstemmed A privacy-preserving solution for compressed storage and selective retrieval of genomic data
title_short A privacy-preserving solution for compressed storage and selective retrieval of genomic data
title_sort privacy-preserving solution for compressed storage and selective retrieval of genomic data
topic Method
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5131820/
https://www.ncbi.nlm.nih.gov/pubmed/27789525
http://dx.doi.org/10.1101/gr.206870.116
work_keys_str_mv AT huangzhicong aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT aydayerman aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT linhuang aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT aiyarraekas aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT molyneauxadam aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT xuzhenyu aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT fellayjacques aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT steinmetzlarsm aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT hubauxjeanpierre aprivacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT huangzhicong privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT aydayerman privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT linhuang privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT aiyarraekas privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT molyneauxadam privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT xuzhenyu privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT fellayjacques privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT steinmetzlarsm privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata
AT hubauxjeanpierre privacypreservingsolutionforcompressedstorageandselectiveretrievalofgenomicdata