Identifying molecular recognition features in intrinsically disordered regions of proteins by transfer learning

Jack Hanson¹, Thomas Litfin², Kuldip Paliwal¹, Yaoqi Zhou²

Affiliations

¹ Signal Processing Laboratory, Griffith University, Brisbane, QLD 4122, Australia.
² Institute for Glycomics, School of Information and Communication Technology, Griffith University, Southport, QLD 4222, Australia.

PMID: 31504193
DOI: 10.1093/bioinformatics/btz691

Identifying molecular recognition features in intrinsically disordered regions of proteins by transfer learning

Jack Hanson et al. Bioinformatics. 2020.

. 2020 Feb 15;36(4):1107-1113.

doi: 10.1093/bioinformatics/btz691.

Authors

Jack Hanson¹, Thomas Litfin², Kuldip Paliwal¹, Yaoqi Zhou²

Affiliations

¹ Signal Processing Laboratory, Griffith University, Brisbane, QLD 4122, Australia.
² Institute for Glycomics, School of Information and Communication Technology, Griffith University, Southport, QLD 4222, Australia.

PMID: 31504193
DOI: 10.1093/bioinformatics/btz691

Abstract

Motivation: Protein intrinsic disorder describes the tendency of sequence residues to not fold into a rigid three-dimensional shape by themselves. However, some of these disordered regions can transition from disorder to order when interacting with another molecule in segments known as molecular recognition features (MoRFs). Previous analysis has shown that these MoRF regions are indirectly encoded within the prediction of residue disorder as low-confidence predictions [i.e. in a semi-disordered state P(D)≈0.5]. Thus, what has been learned for disorder prediction may be transferable to MoRF prediction. Transferring the internal characterization of protein disorder for the prediction of MoRF residues would allow us to take advantage of the large training set available for disorder prediction, enabling the training of larger analytical models than is currently feasible on the small number of currently available annotated MoRF proteins. In this paper, we propose a new method for MoRF prediction by transfer learning from the SPOT-Disorder2 ensemble models built for disorder prediction.

Results: We confirm that directly training on the MoRF set with a randomly initialized model yields substantially poorer performance on independent test sets than by using the transfer-learning-based method SPOT-MoRF, for both deep and simple networks. Its comparison to current state-of-the-art techniques reveals its superior performance in identifying MoRF binding regions in proteins across two independent testing sets, including our new dataset of >800 protein chains. These test chains share <30% sequence similarity to all training and validation proteins used in SPOT-Disorder2 and SPOT-MoRF, and provide a much-needed large-scale update on the performance of current MoRF predictors. The method is expected to be useful in locating functional disordered regions in proteins.

Availability and implementation: SPOT-MoRF and its data are available as a web server and as a standalone program at: http://sparks-lab.org/jack/server/SPOT-MoRF/index.php.

Supplementary information: Supplementary data are available at Bioinformatics online.

PubMed Disclaimer

Cited by

Comparative evaluation of AlphaFold2 and disorder predictors for prediction of intrinsic disorder, disorder content and fully disordered proteins.
Zhao B, Ghadermarzi S, Kurgan L. Zhao B, et al. Comput Struct Biotechnol J. 2023 Jun 2;21:3248-3258. doi: 10.1016/j.csbj.2023.06.001. eCollection 2023. Comput Struct Biotechnol J. 2023. PMID: 38213902 Free PMC article.
RNA secondary structure prediction using an ensemble of two-dimensional deep neural networks and transfer learning.
Singh J, Hanson J, Paliwal K, Zhou Y. Singh J, et al. Nat Commun. 2019 Nov 27;10(1):5407. doi: 10.1038/s41467-019-13395-9. Nat Commun. 2019. PMID: 31776342 Free PMC article.
Computational prediction of disordered binding regions.
Basu S, Kihara D, Kurgan L. Basu S, et al. Comput Struct Biotechnol J. 2023 Feb 10;21:1487-1497. doi: 10.1016/j.csbj.2023.02.018. eCollection 2023. Comput Struct Biotechnol J. 2023. PMID: 36851914 Free PMC article. Review.
Challenges in describing the conformation and dynamics of proteins with ambiguous behavior.
Roca-Martinez J, Lazar T, Gavalda-Garcia J, Bickel D, Pancsa R, Dixit B, Tzavella K, Ramasamy P, Sanchez-Fornaris M, Grau I, Vranken WF. Roca-Martinez J, et al. Front Mol Biosci. 2022 Aug 3;9:959956. doi: 10.3389/fmolb.2022.959956. eCollection 2022. Front Mol Biosci. 2022. PMID: 35992270 Free PMC article.
Intrinsic Disorder and Other Malleable Arsenals of Evolved Protein Multifunctionality.
Aftab A, Sil S, Nath S, Basu A, Basu S. Aftab A, et al. J Mol Evol. 2024 Dec;92(6):669-684. doi: 10.1007/s00239-024-10196-7. Epub 2024 Aug 30. J Mol Evol. 2024. PMID: 39214891 Review.

See all "Cited by" articles

Publication types

Actions

MeSH terms

Actions
Actions
Actions

Substances

Actions

LinkOut - more resources

Full Text Sources
- Ovid Technologies, Inc.
- Silverchair Information Systems
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Identifying molecular recognition features in intrinsically disordered regions of proteins by transfer learning

Affiliations

Identifying molecular recognition features in intrinsically disordered regions of proteins by transfer learning

Authors

Affiliations

Abstract

Similar articles

Cited by

Publication types

MeSH terms

Substances

LinkOut - more resources

Full Text Sources

Miscellaneous