Cargando…

Opfi: A Python package for identifying gene clusters in large genomics and metagenomics data sets

Gene clusters are sets of co-localized, often contiguous genes that together perform specific functions, many of which are relevant to biotechnology. There is a need for software tools that can extract candidate gene clusters from vast amounts of available genomic data. Therefore, we developed Opfi:...

Descripción completa

Detalles Bibliográficos
Autores principales: Hill, Alexis M., Rybarski, James R., Hu, Kuang, Finkelstein, Ilya J., Wilke, Claus O.
Formato: Online Artículo Texto
Lenguaje:English
Publicado: 2021
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9017871/
https://www.ncbi.nlm.nih.gov/pubmed/35445164
http://dx.doi.org/10.21105/joss.03678
Descripción
Sumario:Gene clusters are sets of co-localized, often contiguous genes that together perform specific functions, many of which are relevant to biotechnology. There is a need for software tools that can extract candidate gene clusters from vast amounts of available genomic data. Therefore, we developed Opfi: a modular pipeline for identification of arbitrary gene clusters in assembled genomic or metagenomic sequences. Opfi contains functions for annotation, de-deduplication, and visualization of putative gene clusters. It utilizes a customizable rule-based filtering approach for selection of candidate systems that adhere to user-defined criteria. Opfi is implemented in Python, and is available on the Python Package Index and on Bioconda (Grüning et al., 2018).