Cargando…

Data Exploration and Classification of News Article Reliability: Deep Learning Study

BACKGROUND: During the ongoing COVID-19 pandemic, we are being exposed to large amounts of information each day. This “infodemic” is defined by the World Health Organization as the mass spread of misleading or false information during a pandemic. This spread of misinformation during the infodemic ul...

Descripción completa

Detalles Bibliográficos
Autores principales:	Zhan, Kevin, Li, Yutong, Osmani, Rafay, Wang, Xiaoyu, Cao, Bo
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	JMIR Publications 2022
Materias:	Original Paper
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9516811/ https://www.ncbi.nlm.nih.gov/pubmed/36193330 http://dx.doi.org/10.2196/38839

_version_	1784798785251049472
author	Zhan, Kevin Li, Yutong Osmani, Rafay Wang, Xiaoyu Cao, Bo
author_facet	Zhan, Kevin Li, Yutong Osmani, Rafay Wang, Xiaoyu Cao, Bo
author_sort	Zhan, Kevin
collection	PubMed
description	BACKGROUND: During the ongoing COVID-19 pandemic, we are being exposed to large amounts of information each day. This “infodemic” is defined by the World Health Organization as the mass spread of misleading or false information during a pandemic. This spread of misinformation during the infodemic ultimately leads to misunderstandings of public health orders or direct opposition against public policies. Although there have been efforts to combat misinformation spread, current manual fact-checking methods are insufficient to combat the infodemic. OBJECTIVE: We propose the use of natural language processing (NLP) and machine learning (ML) techniques to build a model that can be used to identify unreliable news articles online. METHODS: First, we preprocessed the ReCOVery data set to obtain 2029 English news articles tagged with COVID-19 keywords from January to May 2020, which are labeled as reliable or unreliable. Data exploration was conducted to determine major differences between reliable and unreliable articles. We built an ensemble deep learning model using the body text, as well as features, such as sentiment, Empath-derived lexical categories, and readability, to classify the reliability. RESULTS: We found that reliable news articles have a higher proportion of neutral sentiment, while unreliable articles have a higher proportion of negative sentiment. Additionally, our analysis demonstrated that reliable articles are easier to read than unreliable articles, in addition to having different lexical categories and keywords. Our new model was evaluated to achieve the following performance metrics: 0.906 area under the curve (AUC), 0.835 specificity, and 0.945 sensitivity. These values are above the baseline performance of the original ReCOVery model. CONCLUSIONS: This paper identified novel differences between reliable and unreliable news articles; moreover, the model was trained using state-of-the-art deep learning techniques. We aim to be able to use our findings to help researchers and the public audience more easily identify false information and unreliable media in their everyday lives.
format	Online Article Text
id	pubmed-9516811
institution	National Center for Biotechnology Information
language	English
publishDate	2022
publisher	JMIR Publications
record_format	MEDLINE/PubMed
spelling	pubmed-95168112022-09-29 Data Exploration and Classification of News Article Reliability: Deep Learning Study Zhan, Kevin Li, Yutong Osmani, Rafay Wang, Xiaoyu Cao, Bo JMIR Infodemiology Original Paper BACKGROUND: During the ongoing COVID-19 pandemic, we are being exposed to large amounts of information each day. This “infodemic” is defined by the World Health Organization as the mass spread of misleading or false information during a pandemic. This spread of misinformation during the infodemic ultimately leads to misunderstandings of public health orders or direct opposition against public policies. Although there have been efforts to combat misinformation spread, current manual fact-checking methods are insufficient to combat the infodemic. OBJECTIVE: We propose the use of natural language processing (NLP) and machine learning (ML) techniques to build a model that can be used to identify unreliable news articles online. METHODS: First, we preprocessed the ReCOVery data set to obtain 2029 English news articles tagged with COVID-19 keywords from January to May 2020, which are labeled as reliable or unreliable. Data exploration was conducted to determine major differences between reliable and unreliable articles. We built an ensemble deep learning model using the body text, as well as features, such as sentiment, Empath-derived lexical categories, and readability, to classify the reliability. RESULTS: We found that reliable news articles have a higher proportion of neutral sentiment, while unreliable articles have a higher proportion of negative sentiment. Additionally, our analysis demonstrated that reliable articles are easier to read than unreliable articles, in addition to having different lexical categories and keywords. Our new model was evaluated to achieve the following performance metrics: 0.906 area under the curve (AUC), 0.835 specificity, and 0.945 sensitivity. These values are above the baseline performance of the original ReCOVery model. CONCLUSIONS: This paper identified novel differences between reliable and unreliable news articles; moreover, the model was trained using state-of-the-art deep learning techniques. We aim to be able to use our findings to help researchers and the public audience more easily identify false information and unreliable media in their everyday lives. JMIR Publications 2022-09-22 /pmc/articles/PMC9516811/ /pubmed/36193330 http://dx.doi.org/10.2196/38839 Text en ©Kevin Zhan, Yutong Li, Rafay Osmani, Xiaoyu Wang, Bo Cao. Originally published in JMIR Infodemiology (https://infodemiology.jmir.org), 22.09.2022. https://creativecommons.org/licenses/by/4.0/This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Infodemiology, is properly cited. The complete bibliographic information, a link to the original publication on https://infodemiology.jmir.org/, as well as this copyright and license information must be included.
spellingShingle	Original Paper Zhan, Kevin Li, Yutong Osmani, Rafay Wang, Xiaoyu Cao, Bo Data Exploration and Classification of News Article Reliability: Deep Learning Study
title	Data Exploration and Classification of News Article Reliability: Deep Learning Study
title_full	Data Exploration and Classification of News Article Reliability: Deep Learning Study
title_fullStr	Data Exploration and Classification of News Article Reliability: Deep Learning Study
title_full_unstemmed	Data Exploration and Classification of News Article Reliability: Deep Learning Study
title_short	Data Exploration and Classification of News Article Reliability: Deep Learning Study
title_sort	data exploration and classification of news article reliability: deep learning study
topic	Original Paper
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9516811/ https://www.ncbi.nlm.nih.gov/pubmed/36193330 http://dx.doi.org/10.2196/38839
work_keys_str_mv	AT zhankevin dataexplorationandclassificationofnewsarticlereliabilitydeeplearningstudy AT liyutong dataexplorationandclassificationofnewsarticlereliabilitydeeplearningstudy AT osmanirafay dataexplorationandclassificationofnewsarticlereliabilitydeeplearningstudy AT wangxiaoyu dataexplorationandclassificationofnewsarticlereliabilitydeeplearningstudy AT caobo dataexplorationandclassificationofnewsarticlereliabilitydeeplearningstudy

Data Exploration and Classification of News Article Reliability: Deep Learning Study

Ejemplares similares