Cargando…

STRIKE: evaluation of protein MSAs using a single 3D structure

Motivation: Evaluating alternative multiple protein sequence alignments is an important unsolved problem in Biology. The most accurate way of doing this is to use structural information. Unfortunately, most methods require at least two structures to be embedded in the alignment, a condition rarely m...

Descripción completa

Detalles Bibliográficos
Autores principales: Kemena, Carsten, Taly, Jean-Francois, Kleinjung, Jens, Notredame, Cedric
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Oxford University Press 2011
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3232373/
https://www.ncbi.nlm.nih.gov/pubmed/22039207
http://dx.doi.org/10.1093/bioinformatics/btr587
Descripción
Sumario:Motivation: Evaluating alternative multiple protein sequence alignments is an important unsolved problem in Biology. The most accurate way of doing this is to use structural information. Unfortunately, most methods require at least two structures to be embedded in the alignment, a condition rarely met when dealing with standard datasets. Result: We developed STRIKE, a method that determines the relative accuracy of two alternative alignments of the same sequences using a single structure. We validated our methodology on three commonly used reference datasets (BAliBASE, Homestrad and Prefab). Given two alignments, STRIKE manages to identify the most accurate one in 70% of the cases on average. This figure increases to 79% when considering very challenging datasets like the RV11 category of BAliBASE. This discrimination capacity is significantly higher than that reported for other metrics such as Contact Accepted mutation or Blosum. We show that this increased performance results both from a refined definition of the contacts and from the use of an improved contact substitution score. Contact: cedric.notredame@crg.eu Availability: STRIKE is an open source freeware available from www.tcoffee.org Supplementary Information: Supplementary data are available at Bioinformatics online.