Cargando…
WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora
In this work we present the design and implementation of WARCProcessor, a novel multiplatform integrative tool aimed to build scientific datasets to facilitate experimentation in web spam research. The developed application allows the user to specify multiple criteria that change the way in which ne...
Autores principales: | , , , , , , |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
MDPI
2017
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5795495/ https://www.ncbi.nlm.nih.gov/pubmed/29271913 http://dx.doi.org/10.3390/s18010016 |
_version_ | 1783297308610265088 |
---|---|
author | Callón, Miguel Fdez-Glez, Jorge Ruano-Ordás, David Laza, Rosalía Pavón, Reyes Fdez-Riverola, Florentino Méndez, Jose Ramón |
author_facet | Callón, Miguel Fdez-Glez, Jorge Ruano-Ordás, David Laza, Rosalía Pavón, Reyes Fdez-Riverola, Florentino Méndez, Jose Ramón |
author_sort | Callón, Miguel |
collection | PubMed |
description | In this work we present the design and implementation of WARCProcessor, a novel multiplatform integrative tool aimed to build scientific datasets to facilitate experimentation in web spam research. The developed application allows the user to specify multiple criteria that change the way in which new corpora are generated whilst reducing the number of repetitive and error prone tasks related with existing corpus maintenance. For this goal, WARCProcessor supports up to six commonly used data sources for web spam research, being able to store output corpus in standard WARC format together with complementary metadata files. Additionally, the application facilitates the automatic and concurrent download of web sites from Internet, giving the possibility of configuring the deep of the links to be followed as well as the behaviour when redirected URLs appear. WARCProcessor supports both an interactive GUI interface and a command line utility for being executed in background. |
format | Online Article Text |
id | pubmed-5795495 |
institution | National Center for Biotechnology Information |
language | English |
publishDate | 2017 |
publisher | MDPI |
record_format | MEDLINE/PubMed |
spelling | pubmed-57954952018-02-13 WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora Callón, Miguel Fdez-Glez, Jorge Ruano-Ordás, David Laza, Rosalía Pavón, Reyes Fdez-Riverola, Florentino Méndez, Jose Ramón Sensors (Basel) Article In this work we present the design and implementation of WARCProcessor, a novel multiplatform integrative tool aimed to build scientific datasets to facilitate experimentation in web spam research. The developed application allows the user to specify multiple criteria that change the way in which new corpora are generated whilst reducing the number of repetitive and error prone tasks related with existing corpus maintenance. For this goal, WARCProcessor supports up to six commonly used data sources for web spam research, being able to store output corpus in standard WARC format together with complementary metadata files. Additionally, the application facilitates the automatic and concurrent download of web sites from Internet, giving the possibility of configuring the deep of the links to be followed as well as the behaviour when redirected URLs appear. WARCProcessor supports both an interactive GUI interface and a command line utility for being executed in background. MDPI 2017-12-22 /pmc/articles/PMC5795495/ /pubmed/29271913 http://dx.doi.org/10.3390/s18010016 Text en © 2017 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/). |
spellingShingle | Article Callón, Miguel Fdez-Glez, Jorge Ruano-Ordás, David Laza, Rosalía Pavón, Reyes Fdez-Riverola, Florentino Méndez, Jose Ramón WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora |
title | WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora |
title_full | WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora |
title_fullStr | WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora |
title_full_unstemmed | WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora |
title_short | WARCProcessor: An Integrative Tool for Building and Management of Web Spam Corpora |
title_sort | warcprocessor: an integrative tool for building and management of web spam corpora |
topic | Article |
url | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5795495/ https://www.ncbi.nlm.nih.gov/pubmed/29271913 http://dx.doi.org/10.3390/s18010016 |
work_keys_str_mv | AT callonmiguel warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora AT fdezglezjorge warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora AT ruanoordasdavid warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora AT lazarosalia warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora AT pavonreyes warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora AT fdezriverolaflorentino warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora AT mendezjoseramon warcprocessoranintegrativetoolforbuildingandmanagementofwebspamcorpora |