Cargando…

metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data

BACKGROUND: Genomic Observatories (GOs) are sites of long-term scientific study that undertake regular assessments of the genomic biodiversity. The European Marine Omics Biodiversity Observation Network (EMO BON) is a network of GOs that conduct regular biological community samplings to generate env...

Descripción completa

Detalles Bibliográficos
Autores principales: Zafeiropoulos, Haris, Beracochea, Martin, Ninidakis, Stelios, Exter, Katrina, Potirakis, Antonis, De Moro, Gianluca, Richardson, Lorna, Corre, Erwan, Machado, João, Pafilis, Evangelos, Kotoulas, Georgios, Santi, Ioulia, Finn, Robert D, Cox, Cymon J, Pavloudi, Christina
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Oxford University Press 2023
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10583283/
https://www.ncbi.nlm.nih.gov/pubmed/37850871
http://dx.doi.org/10.1093/gigascience/giad078
_version_ 1785122516920958976
author Zafeiropoulos, Haris
Beracochea, Martin
Ninidakis, Stelios
Exter, Katrina
Potirakis, Antonis
De Moro, Gianluca
Richardson, Lorna
Corre, Erwan
Machado, João
Pafilis, Evangelos
Kotoulas, Georgios
Santi, Ioulia
Finn, Robert D
Cox, Cymon J
Pavloudi, Christina
author_facet Zafeiropoulos, Haris
Beracochea, Martin
Ninidakis, Stelios
Exter, Katrina
Potirakis, Antonis
De Moro, Gianluca
Richardson, Lorna
Corre, Erwan
Machado, João
Pafilis, Evangelos
Kotoulas, Georgios
Santi, Ioulia
Finn, Robert D
Cox, Cymon J
Pavloudi, Christina
author_sort Zafeiropoulos, Haris
collection PubMed
description BACKGROUND: Genomic Observatories (GOs) are sites of long-term scientific study that undertake regular assessments of the genomic biodiversity. The European Marine Omics Biodiversity Observation Network (EMO BON) is a network of GOs that conduct regular biological community samplings to generate environmental and metagenomic data of microbial communities from designated marine stations around Europe. The development of an effective workflow is essential for the analysis of the EMO BON metagenomic data in a timely and reproducible manner. FINDINGS: Based on the established MGnify resource, we developed metaGOflow. metaGOflow supports the fast inference of taxonomic profiles from GO-derived data based on ribosomal RNA genes and their functional annotation using the raw reads. Thanks to the Research Object Crate packaging, relevant metadata about the sample under study, and the details of the bioinformatics analysis it has been subjected to, are inherited to the data product while its modular implementation allows running the workflow partially. The analysis of 2 EMO BON samples and 1 Tara Oceans sample was performed as a use case. CONCLUSIONS: metaGOflow is an efficient and robust workflow that scales to the needs of projects producing big metagenomic data such as EMO BON. It highlights how containerization technologies along with modern workflow languages and metadata package approaches can support the needs of researchers when dealing with ever-increasing volumes of biological data. Despite being initially oriented to address the needs of EMO BON, metaGOflow is a flexible and easy-to-use workflow that can be broadly used for one-sample-at-a-time analysis of shotgun metagenomics data.
format Online
Article
Text
id pubmed-10583283
institution National Center for Biotechnology Information
language English
publishDate 2023
publisher Oxford University Press
record_format MEDLINE/PubMed
spelling pubmed-105832832023-10-19 metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data Zafeiropoulos, Haris Beracochea, Martin Ninidakis, Stelios Exter, Katrina Potirakis, Antonis De Moro, Gianluca Richardson, Lorna Corre, Erwan Machado, João Pafilis, Evangelos Kotoulas, Georgios Santi, Ioulia Finn, Robert D Cox, Cymon J Pavloudi, Christina Gigascience Technical Note BACKGROUND: Genomic Observatories (GOs) are sites of long-term scientific study that undertake regular assessments of the genomic biodiversity. The European Marine Omics Biodiversity Observation Network (EMO BON) is a network of GOs that conduct regular biological community samplings to generate environmental and metagenomic data of microbial communities from designated marine stations around Europe. The development of an effective workflow is essential for the analysis of the EMO BON metagenomic data in a timely and reproducible manner. FINDINGS: Based on the established MGnify resource, we developed metaGOflow. metaGOflow supports the fast inference of taxonomic profiles from GO-derived data based on ribosomal RNA genes and their functional annotation using the raw reads. Thanks to the Research Object Crate packaging, relevant metadata about the sample under study, and the details of the bioinformatics analysis it has been subjected to, are inherited to the data product while its modular implementation allows running the workflow partially. The analysis of 2 EMO BON samples and 1 Tara Oceans sample was performed as a use case. CONCLUSIONS: metaGOflow is an efficient and robust workflow that scales to the needs of projects producing big metagenomic data such as EMO BON. It highlights how containerization technologies along with modern workflow languages and metadata package approaches can support the needs of researchers when dealing with ever-increasing volumes of biological data. Despite being initially oriented to address the needs of EMO BON, metaGOflow is a flexible and easy-to-use workflow that can be broadly used for one-sample-at-a-time analysis of shotgun metagenomics data. Oxford University Press 2023-10-18 /pmc/articles/PMC10583283/ /pubmed/37850871 http://dx.doi.org/10.1093/gigascience/giad078 Text en © The Author(s) 2023. Published by Oxford University Press GigaScience. https://creativecommons.org/licenses/by/4.0/This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle Technical Note
Zafeiropoulos, Haris
Beracochea, Martin
Ninidakis, Stelios
Exter, Katrina
Potirakis, Antonis
De Moro, Gianluca
Richardson, Lorna
Corre, Erwan
Machado, João
Pafilis, Evangelos
Kotoulas, Georgios
Santi, Ioulia
Finn, Robert D
Cox, Cymon J
Pavloudi, Christina
metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data
title metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data
title_full metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data
title_fullStr metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data
title_full_unstemmed metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data
title_short metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data
title_sort metagoflow: a workflow for the analysis of marine genomic observatories shotgun metagenomics data
topic Technical Note
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10583283/
https://www.ncbi.nlm.nih.gov/pubmed/37850871
http://dx.doi.org/10.1093/gigascience/giad078
work_keys_str_mv AT zafeiropoulosharis metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT beracocheamartin metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT ninidakisstelios metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT exterkatrina metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT potirakisantonis metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT demorogianluca metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT richardsonlorna metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT correerwan metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT machadojoao metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT pafilisevangelos metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT kotoulasgeorgios metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT santiioulia metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT finnrobertd metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT coxcymonj metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata
AT pavloudichristina metagoflowaworkflowfortheanalysisofmarinegenomicobservatoriesshotgunmetagenomicsdata