. 2018 Mar;86 Suppl 1(Suppl 1):67-77.

doi: 10.1002/prot.25377. Epub 2017 Sep 6.

Analysis of deep learning methods for blind protein contact prediction in CASP12

Sheng Wang¹, Siqi Sun¹, Jinbo Xu¹

Affiliations

PMID: 28845538
PMCID: PMC5871922
DOI: 10.1002/prot.25377

Analysis of deep learning methods for blind protein contact prediction in CASP12

Sheng Wang et al. Proteins. 2018 Mar.

. 2018 Mar;86 Suppl 1(Suppl 1):67-77.

doi: 10.1002/prot.25377. Epub 2017 Sep 6.

Authors

Sheng Wang¹, Siqi Sun¹, Jinbo Xu¹

Affiliation

¹ Toyota Technological Institute at Chicago, Chicago, Illinois.

PMID: 28845538
PMCID: PMC5871922
DOI: 10.1002/prot.25377

Abstract

Here we present the results of protein contact prediction achieved in CASP12 by our RaptorX-Contact server, which is an early implementation of our deep learning method for contact prediction. On a set of 38 free-modeling target domains with a median family size of around 58 effective sequences, our server obtained an average top L/5 long- and medium-range contact accuracy of 47% and 44%, respectively (L = length). A complete implementation has an average accuracy of 59% and 57%, respectively. Our deep learning method formulates contact prediction as a pixel-level image labeling problem and simultaneously predicts all residue pairs of a protein using a combination of two deep residual neural networks, taking as input the residue conservation information, predicted secondary structure and solvent accessibility, contact potential, and coevolution information. Our approach differs from existing methods mainly in (1) formulating contact prediction as a pixel-level image labeling problem instead of an image-level classification problem; (2) simultaneously predicting all contacts of an individual protein to make effective use of contact occurrence patterns; and (3) integrating both one-dimensional and two-dimensional deep convolutional neural networks to effectively learn complex sequence-structure relationship including high-order residue correlation. This paper discusses the RaptorX-Contact pipeline, both contact prediction and contact-based folding results, and finally the strength and weakness of our method.

Keywords: CASP; coevolution analysis; deep learning; protein contact prediction; protein folding.

PubMed Disclaimer

Figures

**Figure 1**
(A) The overall network architecture of the deep learning model. Meanwhile, L is protein sequence length and n is the number of hidden neurons in the last 1D convolutional layer. (B) The internal structure of a residual block with X_l and X_l+1 being input and output, respectively.

**Figure 2**
Overlap between predicted contacts (in red and green) and the native (in grey) for T0864-D1. Red (green) dots indicate correct (incorrect) prediction. Top L/2 predicted contacts by each method are shown. (A) The comparison between our prediction (in upper-left triangle) and CCMpred (in lower-right triangle). (B) The comparison between our prediction (in upper-left triangle) and MetaPSICOV-submit (in lower-right triangle).

**Figure 3**
Superimposition between the predicted models (red) and the native structure (blue) for T0864-D1. The models are built by CNS from the contacts predicted by **(A)** our method, **(B)** CCMpred, and **(C)** MetaPSICOV. The TMscores of the three models are 0.63, 0.27 and 0.35, respectively.

**Figure 4**
Overlap between predicted contacts (in red and green) and the native (in grey) for T0869-D1. Red (green) dots indicate correct (incorrect) prediction. Top L/2 predicted contacts by each method are shown. (A) The comparison between our prediction (in upper-left triangle) and CCMpred (in lower-right triangle). (B) The comparison between our prediction (in upper-left triangle) and MetaPSICOV (in lower-right triangle).

**Figure 5**
Superimposition between the predicted models (red) and the native structure (blue) for T0869-D1. The models are built by CNS from the contacts predicted by **(A)** our method, **(B)** CCMpred, and **(C)** MetaPSICOV. Their TMscores are 0.690, 0.265 and 0.441, respectively.

**Figure 6**
Overlap between predicted contacts (in red and green) and the native (in grey). Red (green) dots indicate correct (incorrect) prediction. Top L/2 predicted contacts by each method are shown. (A) The comparison between our prediction (in upper-left triangle) and CCMpred (in lower-right triangle). (B) The comparison between our prediction (in upper-left triangle) and MetaPSICOV (in lower-right triangle).

**Figure 7**
Superimposition between the predicted models (red) and the native structure (blue) for T0904-D1. The models are built by CNS from the contacts predicted by **(A)** our method, **(B)** CCMpred, and **(C)** MetaPSICOV. Their TMscores are 0.682, 0.221 and 0.385, respectively.

See this image and copyright information in PMC

References

1. Kim DE, et al. One contact for every twelve residues allows robust and accurate topology-level protein structure modeling. Proteins. 2014;82(Suppl 2):208–18. - PMC - PubMed
1. de Juan D, Pazos F, Valencia A. Emerging methods in protein co-evolution. Nature reviews Genetics. 2013;14(4):249–61. - PubMed
1. Weigt M, et al. Identification of direct residue contacts in protein-protein interaction by message passing. Proc Natl Acad Sci U S A. 2009;106(1):67–72. - PMC - PubMed
1. Seemayer S, Gruber M, Söding J. CCMpred—fast and precise prediction of protein residue–residue contacts from correlated mutations. Bioinformatics. 2014;30(21):3128–3130. - PMC - PubMed
1. Jones DT, et al. PSICOV: precise structural contact prediction using sparse inverse covariance estimation on large multiple sequence alignments. Bioinformatics. 2012;28(2):184–190. - PubMed

Publication types

Actions
Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions
Actions

Substances

Actions

Grants and funding

R01 GM089753/GM/NIGMS NIH HHS/United States

LinkOut - more resources

Full Text Sources
Other Literature Sources
- scite Smart Citations

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Analysis of deep learning methods for blind protein contact prediction in CASP12

Affiliation

Analysis of deep learning methods for blind protein contact prediction in CASP12

Authors

Affiliation

Abstract

Figures

References

Publication types

MeSH terms

Substances

Grants and funding

LinkOut - more resources

Full Text Sources

Other Literature Sources