Cargando…

Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies

Data harmonization is a key step widely used in multisite neuroimaging studies to remove inter-site heterogeneity of data distribution. However, data harmonization may even introduce additional inter-site differences in neuroimaging data if outliers are present in the data of one or more sites. It r...

Descripción completa

Detalles Bibliográficos
Autores principales: Han, Qichao, Xiao, Xiaoxiao, Wang, Sijia, Qin, Wen, Yu, Chunshui, Liang, Meng
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Frontiers Media S.A. 2023
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10249749/
https://www.ncbi.nlm.nih.gov/pubmed/37304022
http://dx.doi.org/10.3389/fnins.2023.1146175
_version_ 1785055614103191552
author Han, Qichao
Xiao, Xiaoxiao
Wang, Sijia
Qin, Wen
Yu, Chunshui
Liang, Meng
author_facet Han, Qichao
Xiao, Xiaoxiao
Wang, Sijia
Qin, Wen
Yu, Chunshui
Liang, Meng
author_sort Han, Qichao
collection PubMed
description Data harmonization is a key step widely used in multisite neuroimaging studies to remove inter-site heterogeneity of data distribution. However, data harmonization may even introduce additional inter-site differences in neuroimaging data if outliers are present in the data of one or more sites. It remains unclear how the presence of outliers could affect the effectiveness of data harmonization and consequently the results of analyses using harmonized data. To address this question, we generated a normal simulation dataset without outliers and a series of simulation datasets with outliers of varying properties (e.g., outlier location, outlier quantity, and outlier score) based on a real large-sample neuroimaging dataset. We first verified the effectiveness of the most commonly used ComBat harmonization method in the removal of inter-site heterogeneity using the normal simulation data, and then characterized the effects of outliers on the effectiveness of ComBat harmonization and on the results of association analyses between brain imaging-derived phenotypes and a simulated behavioral variable using the simulation datasets with outliers. We found that, although ComBat harmonization effectively removed the inter-site heterogeneity in multisite data and consequently improved the detection of the true brain-behavior relationships, the presence of outliers could damage severely the effectiveness of ComBat harmonization in the removal of data heterogeneity or even introduce extra heterogeneity in the data. Moreover, we found that the effects of outliers on the improvement of the detection of brain-behavior associations by ComBat harmonization were dependent on how such associations were assessed (i.e., by Pearson correlation or Spearman correlation), and on the outlier location, quantity, and outlier score. These findings help us better understand the influences of outliers on data harmonization and highlight the importance of detecting and removing outliers prior to data harmonization in multisite neuroimaging studies.
format Online
Article
Text
id pubmed-10249749
institution National Center for Biotechnology Information
language English
publishDate 2023
publisher Frontiers Media S.A.
record_format MEDLINE/PubMed
spelling pubmed-102497492023-06-09 Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies Han, Qichao Xiao, Xiaoxiao Wang, Sijia Qin, Wen Yu, Chunshui Liang, Meng Front Neurosci Neuroscience Data harmonization is a key step widely used in multisite neuroimaging studies to remove inter-site heterogeneity of data distribution. However, data harmonization may even introduce additional inter-site differences in neuroimaging data if outliers are present in the data of one or more sites. It remains unclear how the presence of outliers could affect the effectiveness of data harmonization and consequently the results of analyses using harmonized data. To address this question, we generated a normal simulation dataset without outliers and a series of simulation datasets with outliers of varying properties (e.g., outlier location, outlier quantity, and outlier score) based on a real large-sample neuroimaging dataset. We first verified the effectiveness of the most commonly used ComBat harmonization method in the removal of inter-site heterogeneity using the normal simulation data, and then characterized the effects of outliers on the effectiveness of ComBat harmonization and on the results of association analyses between brain imaging-derived phenotypes and a simulated behavioral variable using the simulation datasets with outliers. We found that, although ComBat harmonization effectively removed the inter-site heterogeneity in multisite data and consequently improved the detection of the true brain-behavior relationships, the presence of outliers could damage severely the effectiveness of ComBat harmonization in the removal of data heterogeneity or even introduce extra heterogeneity in the data. Moreover, we found that the effects of outliers on the improvement of the detection of brain-behavior associations by ComBat harmonization were dependent on how such associations were assessed (i.e., by Pearson correlation or Spearman correlation), and on the outlier location, quantity, and outlier score. These findings help us better understand the influences of outliers on data harmonization and highlight the importance of detecting and removing outliers prior to data harmonization in multisite neuroimaging studies. Frontiers Media S.A. 2023-05-25 /pmc/articles/PMC10249749/ /pubmed/37304022 http://dx.doi.org/10.3389/fnins.2023.1146175 Text en Copyright © 2023 Han, Xiao, Wang, Qin, Yu and Liang. https://creativecommons.org/licenses/by/4.0/This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
spellingShingle Neuroscience
Han, Qichao
Xiao, Xiaoxiao
Wang, Sijia
Qin, Wen
Yu, Chunshui
Liang, Meng
Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
title Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
title_full Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
title_fullStr Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
title_full_unstemmed Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
title_short Characterization of the effects of outliers on ComBat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
title_sort characterization of the effects of outliers on combat harmonization for removing inter-site data heterogeneity in multisite neuroimaging studies
topic Neuroscience
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10249749/
https://www.ncbi.nlm.nih.gov/pubmed/37304022
http://dx.doi.org/10.3389/fnins.2023.1146175
work_keys_str_mv AT hanqichao characterizationoftheeffectsofoutliersoncombatharmonizationforremovingintersitedataheterogeneityinmultisiteneuroimagingstudies
AT xiaoxiaoxiao characterizationoftheeffectsofoutliersoncombatharmonizationforremovingintersitedataheterogeneityinmultisiteneuroimagingstudies
AT wangsijia characterizationoftheeffectsofoutliersoncombatharmonizationforremovingintersitedataheterogeneityinmultisiteneuroimagingstudies
AT qinwen characterizationoftheeffectsofoutliersoncombatharmonizationforremovingintersitedataheterogeneityinmultisiteneuroimagingstudies
AT yuchunshui characterizationoftheeffectsofoutliersoncombatharmonizationforremovingintersitedataheterogeneityinmultisiteneuroimagingstudies
AT liangmeng characterizationoftheeffectsofoutliersoncombatharmonizationforremovingintersitedataheterogeneityinmultisiteneuroimagingstudies