Cargando…

Detection of changes in literary writing style using N-grams as style markers and supervised machine learning

The analysis of an author’s writing style implies the characterization and identification of the style in terms of a set of features commonly called linguistic features. The analysis can be extrinsic, where the style of an author can be compared with other authors, or intrinsic, where the style of a...

Descripción completa

Detalles Bibliográficos
Autores principales: Ríos-Toledo, Germán, Posadas-Durán, Juan Pablo Francisco, Sidorov, Grigori, Castro-Sánchez, Noé Alejandro
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Public Library of Science 2022
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9299308/
https://www.ncbi.nlm.nih.gov/pubmed/35857768
http://dx.doi.org/10.1371/journal.pone.0267590
_version_ 1784750937978437632
author Ríos-Toledo, Germán
Posadas-Durán, Juan Pablo Francisco
Sidorov, Grigori
Castro-Sánchez, Noé Alejandro
author_facet Ríos-Toledo, Germán
Posadas-Durán, Juan Pablo Francisco
Sidorov, Grigori
Castro-Sánchez, Noé Alejandro
author_sort Ríos-Toledo, Germán
collection PubMed
description The analysis of an author’s writing style implies the characterization and identification of the style in terms of a set of features commonly called linguistic features. The analysis can be extrinsic, where the style of an author can be compared with other authors, or intrinsic, where the style of an author is identified through different stages of his life. Intrinsic analysis has been used, for example, to detect mental illness and the effects of aging. A key element of the analysis is the style markers used to model the author’s writing patterns. The style markers should handle diachronic changes and be thematic independent. One of the most commonly used style marker in extrinsic style analysis is n-gram. In this paper, we present the evaluation of traditional n-grams (words and characters) and dependency tree syntactic n-grams to solve the task of detecting changes in writing style over time. Our corpus consisted of novels by eleven English-speaking authors. The novels of each author were organized chronologically from the oldest to the most recent work according to the date of publication. Subsequently, two stages were defined: initial and final. In each stage three novels were assigned, novels of the initial stage corresponded to the oldest and those at the final stage to the most recent novels. To analyze changes in the writing style, novels were characterized by using four types of n-grams: characters, words, Part-Of-Speech (POS) tags and syntactic relations n-grams. Experiments were performed with a Logistic Regression classifier. Dimension reduction techniques such as Principal Component Analysis (PCA) and Latent Semantic Analysis (LSA) algorithms were evaluated. The results obtained with the different n-grams indicated that all authors presented significant changes in writing style over time. In addition, representations using n-grams of syntactic relations have achieved competitive results among different authors.
format Online
Article
Text
id pubmed-9299308
institution National Center for Biotechnology Information
language English
publishDate 2022
publisher Public Library of Science
record_format MEDLINE/PubMed
spelling pubmed-92993082022-07-21 Detection of changes in literary writing style using N-grams as style markers and supervised machine learning Ríos-Toledo, Germán Posadas-Durán, Juan Pablo Francisco Sidorov, Grigori Castro-Sánchez, Noé Alejandro PLoS One Research Article The analysis of an author’s writing style implies the characterization and identification of the style in terms of a set of features commonly called linguistic features. The analysis can be extrinsic, where the style of an author can be compared with other authors, or intrinsic, where the style of an author is identified through different stages of his life. Intrinsic analysis has been used, for example, to detect mental illness and the effects of aging. A key element of the analysis is the style markers used to model the author’s writing patterns. The style markers should handle diachronic changes and be thematic independent. One of the most commonly used style marker in extrinsic style analysis is n-gram. In this paper, we present the evaluation of traditional n-grams (words and characters) and dependency tree syntactic n-grams to solve the task of detecting changes in writing style over time. Our corpus consisted of novels by eleven English-speaking authors. The novels of each author were organized chronologically from the oldest to the most recent work according to the date of publication. Subsequently, two stages were defined: initial and final. In each stage three novels were assigned, novels of the initial stage corresponded to the oldest and those at the final stage to the most recent novels. To analyze changes in the writing style, novels were characterized by using four types of n-grams: characters, words, Part-Of-Speech (POS) tags and syntactic relations n-grams. Experiments were performed with a Logistic Regression classifier. Dimension reduction techniques such as Principal Component Analysis (PCA) and Latent Semantic Analysis (LSA) algorithms were evaluated. The results obtained with the different n-grams indicated that all authors presented significant changes in writing style over time. In addition, representations using n-grams of syntactic relations have achieved competitive results among different authors. Public Library of Science 2022-07-20 /pmc/articles/PMC9299308/ /pubmed/35857768 http://dx.doi.org/10.1371/journal.pone.0267590 Text en © 2022 Ríos-Toledo et al https://creativecommons.org/licenses/by/4.0/This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/) , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
spellingShingle Research Article
Ríos-Toledo, Germán
Posadas-Durán, Juan Pablo Francisco
Sidorov, Grigori
Castro-Sánchez, Noé Alejandro
Detection of changes in literary writing style using N-grams as style markers and supervised machine learning
title Detection of changes in literary writing style using N-grams as style markers and supervised machine learning
title_full Detection of changes in literary writing style using N-grams as style markers and supervised machine learning
title_fullStr Detection of changes in literary writing style using N-grams as style markers and supervised machine learning
title_full_unstemmed Detection of changes in literary writing style using N-grams as style markers and supervised machine learning
title_short Detection of changes in literary writing style using N-grams as style markers and supervised machine learning
title_sort detection of changes in literary writing style using n-grams as style markers and supervised machine learning
topic Research Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9299308/
https://www.ncbi.nlm.nih.gov/pubmed/35857768
http://dx.doi.org/10.1371/journal.pone.0267590
work_keys_str_mv AT riostoledogerman detectionofchangesinliterarywritingstyleusingngramsasstylemarkersandsupervisedmachinelearning
AT posadasduranjuanpablofrancisco detectionofchangesinliterarywritingstyleusingngramsasstylemarkersandsupervisedmachinelearning
AT sidorovgrigori detectionofchangesinliterarywritingstyleusingngramsasstylemarkersandsupervisedmachinelearning
AT castrosancheznoealejandro detectionofchangesinliterarywritingstyleusingngramsasstylemarkersandsupervisedmachinelearning