Cargando…

Classifying high-dimensional phenotypes with ensemble learning

1. Classification is a fundamental task in biology used to assign members to a class. While linear discriminant functions have long been effective, advances in phenotypic data collection are yielding increasingly high-dimensional datasets with more classes, unequal class covariances, and non-linear...

Descripción completa

Detalles Bibliográficos
Autores principales:	Devine, Jay, Kurki, Helen K., Epp, Jonathan R., Gonzalez, Paula N., Claes, Peter, Hallgrímsson, Benedikt
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Cold Spring Harbor Laboratory 2023
Materias:	Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10312448/ https://www.ncbi.nlm.nih.gov/pubmed/37398168 http://dx.doi.org/10.1101/2023.05.29.542750

_version_	1785066932535296000
author	Devine, Jay Kurki, Helen K. Epp, Jonathan R. Gonzalez, Paula N. Claes, Peter Hallgrímsson, Benedikt
author_facet	Devine, Jay Kurki, Helen K. Epp, Jonathan R. Gonzalez, Paula N. Claes, Peter Hallgrímsson, Benedikt
author_sort	Devine, Jay
collection	PubMed
description	1. Classification is a fundamental task in biology used to assign members to a class. While linear discriminant functions have long been effective, advances in phenotypic data collection are yielding increasingly high-dimensional datasets with more classes, unequal class covariances, and non-linear distributions. Numerous studies have deployed machine learning techniques to classify such distributions, but they are often restricted to a particular organism, a limited set of algorithms, and/or a specific classification task. In addition, the utility of ensemble learning or the strategic combination of models has not been fully explored. 2. We performed a meta-analysis of 33 algorithms across 20 datasets containing over 20,000 high-dimensional shape phenotypes using an ensemble learning framework. Both binary (e.g., sex, environment) and multi-class (e.g., species, genotype, population) classification tasks were considered. The ensemble workflow contains functions for preprocessing, training individual learners and ensembles, and model evaluation. We evaluated algorithm performance within and among datasets. Furthermore, we quantified the extent to which various dataset and phenotypic properties impact performance. 3. We found that discriminant analysis variants and neural networks were the most accurate base learners on average. However, their performance varied substantially between datasets. Ensemble models achieved the highest performance on average, both within and among datasets, increasing average accuracy by up to 3% over the top base learner. Higher class R(2) values, mean class shape distances, and between- vs. within-class variances were positively associated with performance, whereas higher class covariance distances were negatively associated. Class balance and total sample size were not predictive. 4. Learning-based classification is a complex task driven by many hyperparameters. We demonstrate that selecting and optimizing an algorithm based on the results of another study is a flawed strategy. Ensemble models instead offer a flexible approach that is data agnostic and exceptionally accurate. By assessing the impact of various dataset and phenotypic properties on classification performance, we also offer potential explanations for variation in performance. Researchers interested in maximizing performance stand to benefit from the simplicity and effectiveness of our approach made accessible via the R package pheble.
format	Online Article Text
id	pubmed-10312448
institution	National Center for Biotechnology Information
language	English
publishDate	2023
publisher	Cold Spring Harbor Laboratory
record_format	MEDLINE/PubMed
spelling	pubmed-103124482023-07-01 Classifying high-dimensional phenotypes with ensemble learning Devine, Jay Kurki, Helen K. Epp, Jonathan R. Gonzalez, Paula N. Claes, Peter Hallgrímsson, Benedikt bioRxiv Article 1. Classification is a fundamental task in biology used to assign members to a class. While linear discriminant functions have long been effective, advances in phenotypic data collection are yielding increasingly high-dimensional datasets with more classes, unequal class covariances, and non-linear distributions. Numerous studies have deployed machine learning techniques to classify such distributions, but they are often restricted to a particular organism, a limited set of algorithms, and/or a specific classification task. In addition, the utility of ensemble learning or the strategic combination of models has not been fully explored. 2. We performed a meta-analysis of 33 algorithms across 20 datasets containing over 20,000 high-dimensional shape phenotypes using an ensemble learning framework. Both binary (e.g., sex, environment) and multi-class (e.g., species, genotype, population) classification tasks were considered. The ensemble workflow contains functions for preprocessing, training individual learners and ensembles, and model evaluation. We evaluated algorithm performance within and among datasets. Furthermore, we quantified the extent to which various dataset and phenotypic properties impact performance. 3. We found that discriminant analysis variants and neural networks were the most accurate base learners on average. However, their performance varied substantially between datasets. Ensemble models achieved the highest performance on average, both within and among datasets, increasing average accuracy by up to 3% over the top base learner. Higher class R(2) values, mean class shape distances, and between- vs. within-class variances were positively associated with performance, whereas higher class covariance distances were negatively associated. Class balance and total sample size were not predictive. 4. Learning-based classification is a complex task driven by many hyperparameters. We demonstrate that selecting and optimizing an algorithm based on the results of another study is a flawed strategy. Ensemble models instead offer a flexible approach that is data agnostic and exceptionally accurate. By assessing the impact of various dataset and phenotypic properties on classification performance, we also offer potential explanations for variation in performance. Researchers interested in maximizing performance stand to benefit from the simplicity and effectiveness of our approach made accessible via the R package pheble. Cold Spring Harbor Laboratory 2023-05-29 /pmc/articles/PMC10312448/ /pubmed/37398168 http://dx.doi.org/10.1101/2023.05.29.542750 Text en https://creativecommons.org/licenses/by/4.0/This work is licensed under a Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/) , which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use.
spellingShingle	Article Devine, Jay Kurki, Helen K. Epp, Jonathan R. Gonzalez, Paula N. Claes, Peter Hallgrímsson, Benedikt Classifying high-dimensional phenotypes with ensemble learning
title	Classifying high-dimensional phenotypes with ensemble learning
title_full	Classifying high-dimensional phenotypes with ensemble learning
title_fullStr	Classifying high-dimensional phenotypes with ensemble learning
title_full_unstemmed	Classifying high-dimensional phenotypes with ensemble learning
title_short	Classifying high-dimensional phenotypes with ensemble learning
title_sort	classifying high-dimensional phenotypes with ensemble learning
topic	Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10312448/ https://www.ncbi.nlm.nih.gov/pubmed/37398168 http://dx.doi.org/10.1101/2023.05.29.542750
work_keys_str_mv	AT devinejay classifyinghighdimensionalphenotypeswithensemblelearning AT kurkihelenk classifyinghighdimensionalphenotypeswithensemblelearning AT eppjonathanr classifyinghighdimensionalphenotypeswithensemblelearning AT gonzalezpaulan classifyinghighdimensionalphenotypeswithensemblelearning AT claespeter classifyinghighdimensionalphenotypeswithensemblelearning AT hallgrimssonbenedikt classifyinghighdimensionalphenotypeswithensemblelearning

Classifying high-dimensional phenotypes with ensemble learning

Ejemplares similares