Cargando…

Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning

MOTIVATION: T cell heterogeneity presents a challenge for accurate cell identification, understanding their inherent plasticity, and characterizing their critical role in adaptive immunity. Immunologists have traditionally employed techniques such as flow cytometry to identify T cell subtypes based...

Descripción completa

Detalles Bibliográficos
Autores principales: Ran, Ran, Brubaker, Douglas K
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Oxford University Press 2023
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10676521/
https://www.ncbi.nlm.nih.gov/pubmed/38023329
http://dx.doi.org/10.1093/bioadv/vbad159
_version_ 1785141300042924032
author Ran, Ran
Brubaker, Douglas K
author_facet Ran, Ran
Brubaker, Douglas K
author_sort Ran, Ran
collection PubMed
description MOTIVATION: T cell heterogeneity presents a challenge for accurate cell identification, understanding their inherent plasticity, and characterizing their critical role in adaptive immunity. Immunologists have traditionally employed techniques such as flow cytometry to identify T cell subtypes based on a well-established set of surface protein markers. With the advent of single-cell RNA sequencing (scRNA-seq), researchers can now investigate the gene expression profiles of these surface proteins at the single-cell level. The insights gleaned from these profiles offer valuable clues and a deeper understanding of cell identity. However, CD45RA, the isoform of CD45 which distinguishes between naive/central memory T cells and effector memory/effector memory cells re-expressing CD45RA T cells, cannot be well profiled by scRNA-seq due to the difficulty in mapping short reads to genes. RESULTS: In order to facilitate cell-type annotation in T cell scRNA-seq analysis, we employed machine learning and trained a [Formula: see text] classifier on single-cell mRNA count data annotated with known CD45RA antibody levels provided by cellular indexing of transcriptomes and epitopes sequencing data. Among all the algorithms we tested, the trained support vector machine with a radial basis function kernel with optimized hyperparameters achieved a 99.96% accuracy on an unseen dataset. The multilayer perceptron classifier, the second most predictive method overall, also achieved a decent accuracy of 99.74%. Our simple yet robust machine learning approach provides a valid inference on the CD45RA level, assisting the cell identity annotation and further exploring the heterogeneity within human T cells. Based on the overall performance, we chose the support vector machine with a radial basis function kernel as the model implemented in our Python package scCD45RA. AVAILABILITY AND IMPLEMENTATION: The resultant package scCD45RA can be found at https://github.com/BrubakerLab/ScCD45RA and can be installed from the Python Package Index (PyPI) using the command “pip install sccd45ra.”
format Online
Article
Text
id pubmed-10676521
institution National Center for Biotechnology Information
language English
publishDate 2023
publisher Oxford University Press
record_format MEDLINE/PubMed
spelling pubmed-106765212023-11-06 Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning Ran, Ran Brubaker, Douglas K Bioinform Adv Original Article MOTIVATION: T cell heterogeneity presents a challenge for accurate cell identification, understanding their inherent plasticity, and characterizing their critical role in adaptive immunity. Immunologists have traditionally employed techniques such as flow cytometry to identify T cell subtypes based on a well-established set of surface protein markers. With the advent of single-cell RNA sequencing (scRNA-seq), researchers can now investigate the gene expression profiles of these surface proteins at the single-cell level. The insights gleaned from these profiles offer valuable clues and a deeper understanding of cell identity. However, CD45RA, the isoform of CD45 which distinguishes between naive/central memory T cells and effector memory/effector memory cells re-expressing CD45RA T cells, cannot be well profiled by scRNA-seq due to the difficulty in mapping short reads to genes. RESULTS: In order to facilitate cell-type annotation in T cell scRNA-seq analysis, we employed machine learning and trained a [Formula: see text] classifier on single-cell mRNA count data annotated with known CD45RA antibody levels provided by cellular indexing of transcriptomes and epitopes sequencing data. Among all the algorithms we tested, the trained support vector machine with a radial basis function kernel with optimized hyperparameters achieved a 99.96% accuracy on an unseen dataset. The multilayer perceptron classifier, the second most predictive method overall, also achieved a decent accuracy of 99.74%. Our simple yet robust machine learning approach provides a valid inference on the CD45RA level, assisting the cell identity annotation and further exploring the heterogeneity within human T cells. Based on the overall performance, we chose the support vector machine with a radial basis function kernel as the model implemented in our Python package scCD45RA. AVAILABILITY AND IMPLEMENTATION: The resultant package scCD45RA can be found at https://github.com/BrubakerLab/ScCD45RA and can be installed from the Python Package Index (PyPI) using the command “pip install sccd45ra.” Oxford University Press 2023-11-06 /pmc/articles/PMC10676521/ /pubmed/38023329 http://dx.doi.org/10.1093/bioadv/vbad159 Text en © The Author(s) 2023. Published by Oxford University Press. https://creativecommons.org/licenses/by/4.0/This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle Original Article
Ran, Ran
Brubaker, Douglas K
Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning
title Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning
title_full Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning
title_fullStr Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning
title_full_unstemmed Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning
title_short Enhanced annotation of CD45RA to distinguish T cell subsets in single-cell RNA-seq via machine learning
title_sort enhanced annotation of cd45ra to distinguish t cell subsets in single-cell rna-seq via machine learning
topic Original Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10676521/
https://www.ncbi.nlm.nih.gov/pubmed/38023329
http://dx.doi.org/10.1093/bioadv/vbad159
work_keys_str_mv AT ranran enhancedannotationofcd45ratodistinguishtcellsubsetsinsinglecellrnaseqviamachinelearning
AT brubakerdouglask enhancedannotationofcd45ratodistinguishtcellsubsetsinsinglecellrnaseqviamachinelearning