Large-scale automated function prediction of protein sequences and an experimental case study validation on PTEN transcript variants


RİFAİOĞLU A. S. , Dogan T. , Sarac O. S. , Ersahin T., Saidi R., ATALAY M. V. , ...More

PROTEINS-STRUCTURE FUNCTION AND BIOINFORMATICS, vol.86, no.2, pp.135-151, 2018 (Journal Indexed in SCI) identifier identifier identifier

  • Publication Type: Article / Article
  • Volume: 86 Issue: 2
  • Publication Date: 2018
  • Doi Number: 10.1002/prot.25416
  • Title of Journal : PROTEINS-STRUCTURE FUNCTION AND BIOINFORMATICS
  • Page Numbers: pp.135-151
  • Keywords: automated protein function prediction, CHD8, gene ontology, machine learning, protein sequence, PTEN, UniProtKB, variation, FUNCTION ANNOTATION, CLASSIFICATION, SIMILARITY, SEARCH, PFP

Abstract

Recent advances in computing power and machine learning empower functional annotation of protein sequences and their transcript variations. Here, we present an automated prediction system UniGOPred, for GO annotations and a database of GO term predictions for proteomes of several organisms in UniProt Knowledgebase (UniProtKB). UniGOPred provides function predictions for 514 molecular function (MF), 2909 biological process (BP), and 438 cellular component (CC) GO terms for each protein sequence. UniGOPred covers nearly the whole functionality spectrum in Gene Ontology system and it can predict both generic and specific GO terms. UniGOPred was run on CAFA2 challenge target protein sequences and it is categorized within the top 10 best performing methods for the molecular function category. In addition, the performance of UniGOPred is higher compared to the baseline BLAST classifier in all categories of GO. UniGOPred predictions are compared with UniProtKB/TrEMBL database annotations as well. Furthermore, the proposed tool's ability to predict negatively associated GO terms that defines the functions that a protein does not possess, is discussed. UniGOPred annotations were also validated by case studies on PTEN protein variants experimentally and on CHD8 protein variants with literature. UniGOPred protein functional annotation system is available as an open access tool at .