Cargando…

How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking

The lung cancer threat has become a critical issue for public health. Research has been devoted to its clinical study but only a few studies have addressed the issue from a holistic perspective that included social, economic, and environmental dimensions. Therefore, in this study, risk factors or fe...

Descripción completa

Detalles Bibliográficos
Autores principales: Wang, Kung-Min, Chen, Kun-Huang, Hernanda, Chrestella Ayu, Tseng, Shih-Hsien, Wang, Kung-Jeng
Formato: Online Artículo Texto
Lenguaje:English
Publicado: MDPI 2022
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9316771/
https://www.ncbi.nlm.nih.gov/pubmed/35886298
http://dx.doi.org/10.3390/ijerph19148445
_version_ 1784754896451403776
author Wang, Kung-Min
Chen, Kun-Huang
Hernanda, Chrestella Ayu
Tseng, Shih-Hsien
Wang, Kung-Jeng
author_facet Wang, Kung-Min
Chen, Kun-Huang
Hernanda, Chrestella Ayu
Tseng, Shih-Hsien
Wang, Kung-Jeng
author_sort Wang, Kung-Min
collection PubMed
description The lung cancer threat has become a critical issue for public health. Research has been devoted to its clinical study but only a few studies have addressed the issue from a holistic perspective that included social, economic, and environmental dimensions. Therefore, in this study, risk factors or features, such as air pollution, tobacco use, socioeconomic status, employment status, marital status, and environment, were comprehensively considered when constructing a predictive model. These risk factors were analyzed and selected using stepwise regression and the variance inflation factor to eliminate the possibility of multicollinearity. To build efficient and informative prediction models of lung cancer incidence rates, several machine learning algorithms with cross-validation were adopted, namely, linear regression, support vector regression, random forest, K-nearest neighbor, and cubist model tree. A case study in Taiwan showed that the cubist model tree with feature selection was the best model with an RMSE of 3.310 and an R-squared of 0.960. Through these predictive models, we also found that apart from smoking, the average NO(2) concentration, employment percentage, and number of factories were also important factors that had significant impacts on the incidence of lung cancer. In addition, the random forest model without feature selection and with feature selection could support the interpretation of the most contributing variables. The predictive model proposed in the present study can help to precisely analyze and estimate lung cancer incidence rates so that effective preventative measures can be developed. Furthermore, the risk factors involved in the predictive model can help with the future analysis of lung cancer incidence rates from a holistic perspective.
format Online
Article
Text
id pubmed-9316771
institution National Center for Biotechnology Information
language English
publishDate 2022
publisher MDPI
record_format MEDLINE/PubMed
spelling pubmed-93167712022-07-27 How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking Wang, Kung-Min Chen, Kun-Huang Hernanda, Chrestella Ayu Tseng, Shih-Hsien Wang, Kung-Jeng Int J Environ Res Public Health Article The lung cancer threat has become a critical issue for public health. Research has been devoted to its clinical study but only a few studies have addressed the issue from a holistic perspective that included social, economic, and environmental dimensions. Therefore, in this study, risk factors or features, such as air pollution, tobacco use, socioeconomic status, employment status, marital status, and environment, were comprehensively considered when constructing a predictive model. These risk factors were analyzed and selected using stepwise regression and the variance inflation factor to eliminate the possibility of multicollinearity. To build efficient and informative prediction models of lung cancer incidence rates, several machine learning algorithms with cross-validation were adopted, namely, linear regression, support vector regression, random forest, K-nearest neighbor, and cubist model tree. A case study in Taiwan showed that the cubist model tree with feature selection was the best model with an RMSE of 3.310 and an R-squared of 0.960. Through these predictive models, we also found that apart from smoking, the average NO(2) concentration, employment percentage, and number of factories were also important factors that had significant impacts on the incidence of lung cancer. In addition, the random forest model without feature selection and with feature selection could support the interpretation of the most contributing variables. The predictive model proposed in the present study can help to precisely analyze and estimate lung cancer incidence rates so that effective preventative measures can be developed. Furthermore, the risk factors involved in the predictive model can help with the future analysis of lung cancer incidence rates from a holistic perspective. MDPI 2022-07-11 /pmc/articles/PMC9316771/ /pubmed/35886298 http://dx.doi.org/10.3390/ijerph19148445 Text en © 2022 by the authors. https://creativecommons.org/licenses/by/4.0/Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
spellingShingle Article
Wang, Kung-Min
Chen, Kun-Huang
Hernanda, Chrestella Ayu
Tseng, Shih-Hsien
Wang, Kung-Jeng
How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking
title How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking
title_full How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking
title_fullStr How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking
title_full_unstemmed How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking
title_short How Is the Lung Cancer Incidence Rate Associated with Environmental Risks? Machine-Learning-Based Modeling and Benchmarking
title_sort how is the lung cancer incidence rate associated with environmental risks? machine-learning-based modeling and benchmarking
topic Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9316771/
https://www.ncbi.nlm.nih.gov/pubmed/35886298
http://dx.doi.org/10.3390/ijerph19148445
work_keys_str_mv AT wangkungmin howisthelungcancerincidencerateassociatedwithenvironmentalrisksmachinelearningbasedmodelingandbenchmarking
AT chenkunhuang howisthelungcancerincidencerateassociatedwithenvironmentalrisksmachinelearningbasedmodelingandbenchmarking
AT hernandachrestellaayu howisthelungcancerincidencerateassociatedwithenvironmentalrisksmachinelearningbasedmodelingandbenchmarking
AT tsengshihhsien howisthelungcancerincidencerateassociatedwithenvironmentalrisksmachinelearningbasedmodelingandbenchmarking
AT wangkungjeng howisthelungcancerincidencerateassociatedwithenvironmentalrisksmachinelearningbasedmodelingandbenchmarking