Aligning amino acid sequences: comparison of commonly used methods
- PMID: 6100188
- DOI: 10.1007/BF02100085
Aligning amino acid sequences: comparison of commonly used methods
Abstract
We examined two extensive families of protein sequences using four different alignment schemes that employ various degrees of "weighting" in order to determine which approach is most sensitive in establishing relationships. All alignments used a similarity approach based on a general algorithm devised by Needleman and Wunsch. The approaches included a simple program, UM (unitary matrix), whereby only identities are scored; a scheme in which the genetic code is used as a basis for weighting (GC); another that employs a matrix based on structural similarity of amino acids taken together with the genetic basis of mutation (SG); and a fourth that uses the empirical log-odds matrix (LOM) developed by Dayhoff on the basis of observed amino acid replacements. The two sequence families examined were (a) nine different globins and (b) nine different tyrosine kinase-like proteins. It was assumed a priori that all members of a family share common ancestry. In cases where two sequences were more than 30% identical, alignments by all four methods were almost always the same. In cases where the percentage identity was less than 20%, however, there were often significant differences in the alignments. On the average, the Dayhoff LOM approach was the most effective in verifying distant relationships, as judged by an empirical "jumbling test." This was not universally the case, however, and in some instances the simple UM was actually as good or better. Trees constructed on the basis of the various alignments differed with regard to their limb lengths, but had essentially the same branching orders. We suggest some reasons for the different effectivenesses of the four approaches in the two different sequence settings, and offer some rules of thumb for assessing the significance of sequence relationships.
Similar articles
-
Using CLUSTAL for multiple sequence alignments.Methods Enzymol. 1996;266:383-402. doi: 10.1016/s0076-6879(96)66024-8. Methods Enzymol. 1996. PMID: 8743695
-
Progressive sequence alignment as a prerequisite to correct phylogenetic trees.J Mol Evol. 1987;25(4):351-60. doi: 10.1007/BF02603120. J Mol Evol. 1987. PMID: 3118049
-
Profile analysis: detection of distantly related proteins.Proc Natl Acad Sci U S A. 1987 Jul;84(13):4355-8. doi: 10.1073/pnas.84.13.4355. Proc Natl Acad Sci U S A. 1987. PMID: 3474607 Free PMC article.
-
Structural divergence and distant relationships in proteins: evolution of the globins.Curr Opin Struct Biol. 2005 Jun;15(3):290-301. doi: 10.1016/j.sbi.2005.05.008. Curr Opin Struct Biol. 2005. PMID: 15922591 Review.
-
Evolution and taxonomy of positive-strand RNA viruses: implications of comparative analysis of amino acid sequences.Crit Rev Biochem Mol Biol. 1993;28(5):375-430. doi: 10.3109/10409239309078440. Crit Rev Biochem Mol Biol. 1993. PMID: 8269709 Review.
Cited by
-
The evolution of rhodopsins and neurotransmitter receptors.J Mol Evol. 1991 Oct;33(4):367-78. doi: 10.1007/BF02102867. J Mol Evol. 1991. PMID: 1663559
-
Sequence similarity between Borna disease virus p40 and a duplicated domain within the paramyxovirus and rhabdovirus polymerase proteins.J Virol. 1992 Nov;66(11):6572-7. doi: 10.1128/JVI.66.11.6572-6577.1992. J Virol. 1992. PMID: 1404604 Free PMC article.
-
Evidence that a human soluble beta-galactoside-binding lectin is encoded by a family of genes.Proc Natl Acad Sci U S A. 1986 Oct;83(20):7603-7. doi: 10.1073/pnas.83.20.7603. Proc Natl Acad Sci U S A. 1986. PMID: 3020551 Free PMC article.
-
Homing endonucleases encoded by germ line-limited genes in Tetrahymena thermophila have APETELA2 DNA binding domains.Eukaryot Cell. 2004 Jun;3(3):685-94. doi: 10.1128/EC.3.3.685-694.2004. Eukaryot Cell. 2004. PMID: 15189989 Free PMC article.
-
Larger rearranged mitochondrial genomes in Dekkera/Brettanomyces yeasts are more closely related than smaller genomes with a conserved gene order.J Mol Evol. 1993 Mar;36(3):263-9. doi: 10.1007/BF00160482. J Mol Evol. 1993. PMID: 8387113
References
Publication types
MeSH terms
Substances
Grants and funding
LinkOut - more resources
Other Literature Sources
Miscellaneous