Cargando…

Soft Bigram distance for names matching

BACKGROUND: Bi-gram distance (BI-DIST) is a recent approach to measure the distance between two strings that have an important role in a wide range of applications in various areas. The importance of BI-DIST is due to its representational and computational efficiency, which has led to extensive rese...

Descripción completa

Detalles Bibliográficos
Autores principales: Hadwan, Mohammed, Al-Hagery, Mohammed A., Al-Sanabani, Maher, Al-Hagree, Salah
Formato: Online Artículo Texto
Lenguaje:English
Publicado: PeerJ Inc. 2021
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8080420/
https://www.ncbi.nlm.nih.gov/pubmed/33981836
http://dx.doi.org/10.7717/peerj-cs.465
Descripción
Sumario:BACKGROUND: Bi-gram distance (BI-DIST) is a recent approach to measure the distance between two strings that have an important role in a wide range of applications in various areas. The importance of BI-DIST is due to its representational and computational efficiency, which has led to extensive research to further enhance its efficiency. However, developing an algorithm that can measure the distance of strings accurately and efficiently has posed a major challenge to many developers. Consequently, this research aims to design an algorithm that can match the names accurately. BI-DIST distance is considered the best orthographic measure for names identification; nevertheless, it lacks a distance scale between the name bigrams. METHODS: In this research, the Soft Bigram Distance (Soft-Bidist) measure is proposed. It is an extension of BI-DIST by softening the scale of comparison among the name Bigrams for improving the name matching. Different datasets are used to demonstrate the efficiency of the proposed method. RESULTS: The results show that Soft-Bidist outperforms the compared algorithms using different name matching datasets.