Cargando…
Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer
BACKGROUND: De novo inference of clinically relevant gene function relationships from tumor RNA-seq remains a challenging task. Current methods typically either partition patient samples into a few subtypes or rely upon analysis of pairwise gene correlations that will miss some groups in noisy data....
Autores principales: | , |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
BioMed Central
2017
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5351169/ https://www.ncbi.nlm.nih.gov/pubmed/28292312 http://dx.doi.org/10.1186/s12920-017-0245-6 |
_version_ | 1782514720972996608 |
---|---|
author | Pepke, Shirley Ver Steeg, Greg |
author_facet | Pepke, Shirley Ver Steeg, Greg |
author_sort | Pepke, Shirley |
collection | PubMed |
description | BACKGROUND: De novo inference of clinically relevant gene function relationships from tumor RNA-seq remains a challenging task. Current methods typically either partition patient samples into a few subtypes or rely upon analysis of pairwise gene correlations that will miss some groups in noisy data. Leveraging higher dimensional information can be expected to increase the power to discern targetable pathways, but this is commonly thought to be an intractable computational problem. METHODS: In this work we adapt a recently developed machine learning algorithm for sensitive detection of complex gene relationships. The algorithm, CorEx, efficiently optimizes over multivariate mutual information and can be iteratively applied to generate a hierarchy of relatively independent latent factors. The learned latent factors are used to stratify patients for survival analysis with respect to both single factors and combinations. These analyses are performed and interpreted in the context of biological function annotations and protein network interactions that might be utilized to match patients to multiple therapies. RESULTS: Analysis of ovarian tumor RNA-seq samples demonstrates the algorithm’s power to infer well over one hundred biologically interpretable gene cohorts, several times more than standard methods such as hierarchical clustering and k-means. The CorEx factor hierarchy is also informative, with related but distinct gene clusters grouped by upper nodes. Some latent factors correlate with patient survival, including one for a pathway connected with the epithelial-mesenchymal transition in breast cancer that is regulated by a microRNA that modulates epigenetics. Further, combinations of factors lead to a synergistic survival advantage in some cases. CONCLUSIONS: In contrast to studies that attempt to partition patients into a small number of subtypes (typically 4 or fewer) for treatment purposes, our approach utilizes subgroup information for combinatoric transcriptional phenotyping. Considering only the 66 gene expression groups that are found to both have significant Gene Ontology enrichment and are small enough to indicate specific drug targets implies a computational phenotype for ovarian cancer that allows for 3(66) possible patient profiles, enabling truly personalized treatment. The findings here demonstrate a new technique that sheds light on the complexity of gene expression dependencies in tumors and could eventually enable the use of patient RNA-seq profiles for selection of personalized and effective cancer treatments. ELECTRONIC SUPPLEMENTARY MATERIAL: The online version of this article (doi:10.1186/s12920-017-0245-6) contains supplementary material, which is available to authorized users. |
format | Online Article Text |
id | pubmed-5351169 |
institution | National Center for Biotechnology Information |
language | English |
publishDate | 2017 |
publisher | BioMed Central |
record_format | MEDLINE/PubMed |
spelling | pubmed-53511692017-03-17 Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer Pepke, Shirley Ver Steeg, Greg BMC Med Genomics Research Article BACKGROUND: De novo inference of clinically relevant gene function relationships from tumor RNA-seq remains a challenging task. Current methods typically either partition patient samples into a few subtypes or rely upon analysis of pairwise gene correlations that will miss some groups in noisy data. Leveraging higher dimensional information can be expected to increase the power to discern targetable pathways, but this is commonly thought to be an intractable computational problem. METHODS: In this work we adapt a recently developed machine learning algorithm for sensitive detection of complex gene relationships. The algorithm, CorEx, efficiently optimizes over multivariate mutual information and can be iteratively applied to generate a hierarchy of relatively independent latent factors. The learned latent factors are used to stratify patients for survival analysis with respect to both single factors and combinations. These analyses are performed and interpreted in the context of biological function annotations and protein network interactions that might be utilized to match patients to multiple therapies. RESULTS: Analysis of ovarian tumor RNA-seq samples demonstrates the algorithm’s power to infer well over one hundred biologically interpretable gene cohorts, several times more than standard methods such as hierarchical clustering and k-means. The CorEx factor hierarchy is also informative, with related but distinct gene clusters grouped by upper nodes. Some latent factors correlate with patient survival, including one for a pathway connected with the epithelial-mesenchymal transition in breast cancer that is regulated by a microRNA that modulates epigenetics. Further, combinations of factors lead to a synergistic survival advantage in some cases. CONCLUSIONS: In contrast to studies that attempt to partition patients into a small number of subtypes (typically 4 or fewer) for treatment purposes, our approach utilizes subgroup information for combinatoric transcriptional phenotyping. Considering only the 66 gene expression groups that are found to both have significant Gene Ontology enrichment and are small enough to indicate specific drug targets implies a computational phenotype for ovarian cancer that allows for 3(66) possible patient profiles, enabling truly personalized treatment. The findings here demonstrate a new technique that sheds light on the complexity of gene expression dependencies in tumors and could eventually enable the use of patient RNA-seq profiles for selection of personalized and effective cancer treatments. ELECTRONIC SUPPLEMENTARY MATERIAL: The online version of this article (doi:10.1186/s12920-017-0245-6) contains supplementary material, which is available to authorized users. BioMed Central 2017-03-15 /pmc/articles/PMC5351169/ /pubmed/28292312 http://dx.doi.org/10.1186/s12920-017-0245-6 Text en © The Author(s) 2017 Open Access This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made. The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in this article, unless otherwise stated. |
spellingShingle | Research Article Pepke, Shirley Ver Steeg, Greg Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
title | Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
title_full | Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
title_fullStr | Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
title_full_unstemmed | Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
title_short | Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
title_sort | comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer |
topic | Research Article |
url | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5351169/ https://www.ncbi.nlm.nih.gov/pubmed/28292312 http://dx.doi.org/10.1186/s12920-017-0245-6 |
work_keys_str_mv | AT pepkeshirley comprehensivediscoveryofsubsamplegeneexpressioncomponentsbyinformationexplanationtherapeuticimplicationsincancer AT versteeggreg comprehensivediscoveryofsubsamplegeneexpressioncomponentsbyinformationexplanationtherapeuticimplicationsincancer |