Cargando…

PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods

MOTIVATION: Novel machine learning and statistical modeling studies rely on standardized comparisons to existing methods using well-studied benchmark datasets. Few tools exist that provide rapid access to many of these datasets through a standardized, user-friendly interface that integrates well wit...

Descripción completa

Detalles Bibliográficos
Autores principales:	Romano, Joseph D, Le, Trang T, La Cava, William, Gregg, John T, Goldberg, Daniel J, Chakraborty, Praneel, Ray, Natasha L, Himmelstein, Daniel, Fu, Weixuan, Moore, Jason H
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Oxford University Press 2021
Materias:	Applications Notes
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8756190/ https://www.ncbi.nlm.nih.gov/pubmed/34677586 http://dx.doi.org/10.1093/bioinformatics/btab727

_version_	1784632515257958400
author	Romano, Joseph D Le, Trang T La Cava, William Gregg, John T Goldberg, Daniel J Chakraborty, Praneel Ray, Natasha L Himmelstein, Daniel Fu, Weixuan Moore, Jason H
author_facet	Romano, Joseph D Le, Trang T La Cava, William Gregg, John T Goldberg, Daniel J Chakraborty, Praneel Ray, Natasha L Himmelstein, Daniel Fu, Weixuan Moore, Jason H
author_sort	Romano, Joseph D
collection	PubMed
description	MOTIVATION: Novel machine learning and statistical modeling studies rely on standardized comparisons to existing methods using well-studied benchmark datasets. Few tools exist that provide rapid access to many of these datasets through a standardized, user-friendly interface that integrates well with popular data science workflows. RESULTS: This release of PMLB (Penn Machine Learning Benchmarks) provides the largest collection of diverse, public benchmark datasets for evaluating new machine learning and data science methods aggregated in one location. v1.0 introduces a number of critical improvements developed following discussions with the open-source community. AVAILABILITY AND IMPLEMENTATION: PMLB is available at https://github.com/EpistasisLab/pmlb. Python and R interfaces for PMLB can be installed through the Python Package Index and Comprehensive R Archive Network, respectively.
format	Online Article Text
id	pubmed-8756190
institution	National Center for Biotechnology Information
language	English
publishDate	2021
publisher	Oxford University Press
record_format	MEDLINE/PubMed
spelling	pubmed-87561902022-01-13 PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods Romano, Joseph D Le, Trang T La Cava, William Gregg, John T Goldberg, Daniel J Chakraborty, Praneel Ray, Natasha L Himmelstein, Daniel Fu, Weixuan Moore, Jason H Bioinformatics Applications Notes MOTIVATION: Novel machine learning and statistical modeling studies rely on standardized comparisons to existing methods using well-studied benchmark datasets. Few tools exist that provide rapid access to many of these datasets through a standardized, user-friendly interface that integrates well with popular data science workflows. RESULTS: This release of PMLB (Penn Machine Learning Benchmarks) provides the largest collection of diverse, public benchmark datasets for evaluating new machine learning and data science methods aggregated in one location. v1.0 introduces a number of critical improvements developed following discussions with the open-source community. AVAILABILITY AND IMPLEMENTATION: PMLB is available at https://github.com/EpistasisLab/pmlb. Python and R interfaces for PMLB can be installed through the Python Package Index and Comprehensive R Archive Network, respectively. Oxford University Press 2021-10-22 /pmc/articles/PMC8756190/ /pubmed/34677586 http://dx.doi.org/10.1093/bioinformatics/btab727 Text en © The Author(s) 2021. Published by Oxford University Press. https://creativecommons.org/licenses/by/4.0/This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle	Applications Notes Romano, Joseph D Le, Trang T La Cava, William Gregg, John T Goldberg, Daniel J Chakraborty, Praneel Ray, Natasha L Himmelstein, Daniel Fu, Weixuan Moore, Jason H PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
title	PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
title_full	PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
title_fullStr	PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
title_full_unstemmed	PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
title_short	PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
title_sort	pmlb v1.0: an open-source dataset collection for benchmarking machine learning methods
topic	Applications Notes
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8756190/ https://www.ncbi.nlm.nih.gov/pubmed/34677586 http://dx.doi.org/10.1093/bioinformatics/btab727
work_keys_str_mv	AT romanojosephd pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT letrangt pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT lacavawilliam pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT greggjohnt pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT goldbergdanielj pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT chakrabortypraneel pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT raynatashal pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT himmelsteindaniel pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT fuweixuan pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods AT moorejasonh pmlbv10anopensourcedatasetcollectionforbenchmarkingmachinelearningmethods

PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods

Ejemplares similares