Cargando…

PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences

SIMPLE SUMMARY: The family of coronaviruses comprises a diverse set of strains and variants which cause diseases from the common cold to COVID-19. Moreover, they infect a wide array of hosts from bats, camels, birds, to humans. Studying coronaviruses through the lens of host specificity provides a u...

Descripción completa

Detalles Bibliográficos
Autores principales: Ali, Sarwan, Bello, Babatunde, Chourasia, Prakash, Punathil, Ria Thazhe, Zhou, Yijing, Patterson, Murray
Formato: Online Artículo Texto
Lenguaje:English
Publicado: MDPI 2022
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8945605/
https://www.ncbi.nlm.nih.gov/pubmed/35336792
http://dx.doi.org/10.3390/biology11030418
_version_ 1784673993757818880
author Ali, Sarwan
Bello, Babatunde
Chourasia, Prakash
Punathil, Ria Thazhe
Zhou, Yijing
Patterson, Murray
author_facet Ali, Sarwan
Bello, Babatunde
Chourasia, Prakash
Punathil, Ria Thazhe
Zhou, Yijing
Patterson, Murray
author_sort Ali, Sarwan
collection PubMed
description SIMPLE SUMMARY: The family of coronaviruses comprises a diverse set of strains and variants which cause diseases from the common cold to COVID-19. Moreover, they infect a wide array of hosts from bats, camels, birds, to humans. Studying coronaviruses through the lens of host specificity provides a unique perspective to understanding the evolution, diversity and dynamics of this family. In particular, this can reveal groups of different hosts infected by similar strains, giving clues on strains which were more likely to have evolved to jump from one host to another. In this work, we frame host specificity as a classification task, in designing a very compact numerical representation of the spike sequences of different coronaviruses. Based on this numerical representation, classification methods are able to detect the target host with high accuracy. Such an approach can used to efficiently scale to large volumes of sequences, in order to unveil trends in the host specificity of different coronavirus strains. ABSTRACT: The study of host specificity has important connections to the question about the origin of SARS-CoV-2 in humans which led to the COVID-19 pandemic—an important open question. There are speculations that bats are a possible origin. Likewise, there are many closely related (corona)viruses, such as SARS, which was found to be transmitted through civets. The study of the different hosts which can be potential carriers and transmitters of deadly viruses to humans is crucial to understanding, mitigating, and preventing current and future pandemics. In coronaviruses, the surface (S) protein, or spike protein, is important in determining host specificity, since it is the point of contact between the virus and the host cell membrane. In this paper, we classify the hosts of over five thousand coronaviruses from their spike protein sequences, segregating them into clusters of distinct hosts among birds, bats, camels, swine, humans, and weasels, to name a few. We propose a feature embedding based on the well-known position weight matrix (PWM), which we call PWM2Vec, and we use it to generate feature vectors from the spike protein sequences of these coronaviruses. While our embedding is inspired by the success of PWMs in biological applications, such as determining protein function and identifying transcription factor binding sites, we are the first (to the best of our knowledge) to use PWMs from viral sequences to generate fixed-length feature vector representations, and use them in the context of host classification. The results on real world data show that when using PWM2Vec, machine learning classifiers are able to perform comparably to the baseline models in terms of predictive performance and runtime—in some cases, the performance is better. We also measure the importance of different amino acids using information gain to show the amino acids which are important for predicting the host of a given coronavirus. Finally, we perform some statistical analyses on these results to show that our embedding is more compact than the embeddings of the baseline models.
format Online
Article
Text
id pubmed-8945605
institution National Center for Biotechnology Information
language English
publishDate 2022
publisher MDPI
record_format MEDLINE/PubMed
spelling pubmed-89456052022-03-25 PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences Ali, Sarwan Bello, Babatunde Chourasia, Prakash Punathil, Ria Thazhe Zhou, Yijing Patterson, Murray Biology (Basel) Article SIMPLE SUMMARY: The family of coronaviruses comprises a diverse set of strains and variants which cause diseases from the common cold to COVID-19. Moreover, they infect a wide array of hosts from bats, camels, birds, to humans. Studying coronaviruses through the lens of host specificity provides a unique perspective to understanding the evolution, diversity and dynamics of this family. In particular, this can reveal groups of different hosts infected by similar strains, giving clues on strains which were more likely to have evolved to jump from one host to another. In this work, we frame host specificity as a classification task, in designing a very compact numerical representation of the spike sequences of different coronaviruses. Based on this numerical representation, classification methods are able to detect the target host with high accuracy. Such an approach can used to efficiently scale to large volumes of sequences, in order to unveil trends in the host specificity of different coronavirus strains. ABSTRACT: The study of host specificity has important connections to the question about the origin of SARS-CoV-2 in humans which led to the COVID-19 pandemic—an important open question. There are speculations that bats are a possible origin. Likewise, there are many closely related (corona)viruses, such as SARS, which was found to be transmitted through civets. The study of the different hosts which can be potential carriers and transmitters of deadly viruses to humans is crucial to understanding, mitigating, and preventing current and future pandemics. In coronaviruses, the surface (S) protein, or spike protein, is important in determining host specificity, since it is the point of contact between the virus and the host cell membrane. In this paper, we classify the hosts of over five thousand coronaviruses from their spike protein sequences, segregating them into clusters of distinct hosts among birds, bats, camels, swine, humans, and weasels, to name a few. We propose a feature embedding based on the well-known position weight matrix (PWM), which we call PWM2Vec, and we use it to generate feature vectors from the spike protein sequences of these coronaviruses. While our embedding is inspired by the success of PWMs in biological applications, such as determining protein function and identifying transcription factor binding sites, we are the first (to the best of our knowledge) to use PWMs from viral sequences to generate fixed-length feature vector representations, and use them in the context of host classification. The results on real world data show that when using PWM2Vec, machine learning classifiers are able to perform comparably to the baseline models in terms of predictive performance and runtime—in some cases, the performance is better. We also measure the importance of different amino acids using information gain to show the amino acids which are important for predicting the host of a given coronavirus. Finally, we perform some statistical analyses on these results to show that our embedding is more compact than the embeddings of the baseline models. MDPI 2022-03-09 /pmc/articles/PMC8945605/ /pubmed/35336792 http://dx.doi.org/10.3390/biology11030418 Text en © 2022 by the authors. https://creativecommons.org/licenses/by/4.0/Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
spellingShingle Article
Ali, Sarwan
Bello, Babatunde
Chourasia, Prakash
Punathil, Ria Thazhe
Zhou, Yijing
Patterson, Murray
PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences
title PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences
title_full PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences
title_fullStr PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences
title_full_unstemmed PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences
title_short PWM2Vec: An Efficient Embedding Approach for Viral Host Specification from Coronavirus Spike Sequences
title_sort pwm2vec: an efficient embedding approach for viral host specification from coronavirus spike sequences
topic Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8945605/
https://www.ncbi.nlm.nih.gov/pubmed/35336792
http://dx.doi.org/10.3390/biology11030418
work_keys_str_mv AT alisarwan pwm2vecanefficientembeddingapproachforviralhostspecificationfromcoronavirusspikesequences
AT bellobabatunde pwm2vecanefficientembeddingapproachforviralhostspecificationfromcoronavirusspikesequences
AT chourasiaprakash pwm2vecanefficientembeddingapproachforviralhostspecificationfromcoronavirusspikesequences
AT punathilriathazhe pwm2vecanefficientembeddingapproachforviralhostspecificationfromcoronavirusspikesequences
AT zhouyijing pwm2vecanefficientembeddingapproachforviralhostspecificationfromcoronavirusspikesequences
AT pattersonmurray pwm2vecanefficientembeddingapproachforviralhostspecificationfromcoronavirusspikesequences