Cargando…
Probabilistic variable-length segmentation of protein sequences for discriminative motif discovery (DiMotif) and sequence embedding (ProtVecX)
In this paper, we present peptide-pair encoding (PPE), a general-purpose probabilistic segmentation of protein sequences into commonly occurring variable-length sub-sequences. The idea of PPE segmentation is inspired by the byte-pair encoding (BPE) text compression algorithm, which has recently gain...
Autores principales: | Asgari, Ehsaneddin, McHardy, Alice C., Mofrad, Mohammad R. K. |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
Nature Publishing Group UK
2019
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6401088/ https://www.ncbi.nlm.nih.gov/pubmed/30837494 http://dx.doi.org/10.1038/s41598-019-38746-w |
Ejemplares similares
-
EpitopeVec: linear epitope prediction using deep protein sequence embeddings
por: Bahai, Akash, et al.
Publicado: (2021) -
MicroPheno: predicting environments and host phenotypes from 16S rRNA gene sequencing using a k-mer based representation of shallow sub-samples
por: Asgari, Ehsaneddin, et al.
Publicado: (2018) -
MicroPheno: predicting environments and host phenotypes from 16S rRNA gene sequencing using a k-mer based representation of shallow sub-samples
por: Asgari, Ehsaneddin, et al.
Publicado: (2019) -
Continuous Distributed Representation of Biological Sequences for Deep Proteomics and Genomics
por: Asgari, Ehsaneddin, et al.
Publicado: (2015) -
GO Bench: shared hub for universal benchmarking of machine learning-based protein functional annotations
por: Dickson, Andrew, et al.
Publicado: (2023)