Cargando…

Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes

BACKGROUND: Artificial intelligence has an emerging progress in diagnostic pathology. A large number of studies of applying deep learning models to histopathological images have been published in recent years. While many studies claim high accuracies, they may fall into the pitfalls of overfitting a...

Descripción completa

Detalles Bibliográficos
Autores principales:	Tang, Haiming, Sun, Nanfei, Shen, Steven
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Wolters Kluwer - Medknow 2021
Materias:	Original Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8404558/ https://www.ncbi.nlm.nih.gov/pubmed/34497734 http://dx.doi.org/10.4103/jpi.jpi_78_20

_version_	1783746189556973568
author	Tang, Haiming Sun, Nanfei Shen, Steven
author_facet	Tang, Haiming Sun, Nanfei Shen, Steven
author_sort	Tang, Haiming
collection	PubMed
description	BACKGROUND: Artificial intelligence has an emerging progress in diagnostic pathology. A large number of studies of applying deep learning models to histopathological images have been published in recent years. While many studies claim high accuracies, they may fall into the pitfalls of overfitting and lack of generalization due to the high variability of the histopathological images. AIMS AND OBJECTS: Use the model training of osteosarcoma as an example to illustrate the pitfalls of overfitting and how the addition of model input variability can help improve model performance. MATERIALS AND METHODS: We use the publicly available osteosarcoma dataset to retrain a previously published classification model for osteosarcoma. We partition the same set of images into the training and testing datasets differently than the original study: the test dataset consists of images from one patient while the training dataset consists images of all other patients. We also show the influence of training data variability on model performance by collecting a minimal dataset of 10 osteosarcoma subtypes as well as benign tissues and benign bone tumors of differentiation. RESULTS: The performance of the re-trained model on the test set using the new partition schema declines dramatically, indicating a lack of model generalization and overfitting. We show the additions of more and moresubtypes into the training data step by step under the same model schema yield a series of coherent models with increasing performances. CONCLUSIONS: In conclusion, we bring forward data preprocessing and collection tactics for histopathological images of high variability to avoid the pitfalls of overfitting and build deep learning models of higher generalization abilities.
format	Online Article Text
id	pubmed-8404558
institution	National Center for Biotechnology Information
language	English
publishDate	2021
publisher	Wolters Kluwer - Medknow
record_format	MEDLINE/PubMed
spelling	pubmed-84045582021-09-07 Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes Tang, Haiming Sun, Nanfei Shen, Steven J Pathol Inform Original Article BACKGROUND: Artificial intelligence has an emerging progress in diagnostic pathology. A large number of studies of applying deep learning models to histopathological images have been published in recent years. While many studies claim high accuracies, they may fall into the pitfalls of overfitting and lack of generalization due to the high variability of the histopathological images. AIMS AND OBJECTS: Use the model training of osteosarcoma as an example to illustrate the pitfalls of overfitting and how the addition of model input variability can help improve model performance. MATERIALS AND METHODS: We use the publicly available osteosarcoma dataset to retrain a previously published classification model for osteosarcoma. We partition the same set of images into the training and testing datasets differently than the original study: the test dataset consists of images from one patient while the training dataset consists images of all other patients. We also show the influence of training data variability on model performance by collecting a minimal dataset of 10 osteosarcoma subtypes as well as benign tissues and benign bone tumors of differentiation. RESULTS: The performance of the re-trained model on the test set using the new partition schema declines dramatically, indicating a lack of model generalization and overfitting. We show the additions of more and moresubtypes into the training data step by step under the same model schema yield a series of coherent models with increasing performances. CONCLUSIONS: In conclusion, we bring forward data preprocessing and collection tactics for histopathological images of high variability to avoid the pitfalls of overfitting and build deep learning models of higher generalization abilities. Wolters Kluwer - Medknow 2021-08-04 /pmc/articles/PMC8404558/ /pubmed/34497734 http://dx.doi.org/10.4103/jpi.jpi_78_20 Text en Copyright: © 2021 Journal of Pathology Informatics https://creativecommons.org/licenses/by-nc-sa/4.0/This is an open access journal, and articles are distributed under the terms of the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License, which allows others to remix, tweak, and build upon the work non-commercially, as long as appropriate credit is given and the new creations are licensed under the identical terms.
spellingShingle	Original Article Tang, Haiming Sun, Nanfei Shen, Steven Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes
title	Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes
title_full	Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes
title_fullStr	Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes
title_full_unstemmed	Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes
title_short	Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes
title_sort	improving generalization of deep learning models for diagnostic pathology by increasing variability in training data: experiments on osteosarcoma subtypes
topic	Original Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8404558/ https://www.ncbi.nlm.nih.gov/pubmed/34497734 http://dx.doi.org/10.4103/jpi.jpi_78_20
work_keys_str_mv	AT tanghaiming improvinggeneralizationofdeeplearningmodelsfordiagnosticpathologybyincreasingvariabilityintrainingdataexperimentsonosteosarcomasubtypes AT sunnanfei improvinggeneralizationofdeeplearningmodelsfordiagnosticpathologybyincreasingvariabilityintrainingdataexperimentsonosteosarcomasubtypes AT shensteven improvinggeneralizationofdeeplearningmodelsfordiagnosticpathologybyincreasingvariabilityintrainingdataexperimentsonosteosarcomasubtypes

Improving Generalization of Deep Learning Models for Diagnostic Pathology by Increasing Variability in Training Data: Experiments on Osteosarcoma Subtypes

Ejemplares similares