UPCLASS: a Deep Learning-based Classifier for UniProtKB Entry Publications
UPCLASS: a Deep Learning-based Classifier for UniProtKB Entry Publications
AbstractIn the UniProt Knowledgebase (UniProtKB), publications providing evidence for a specific protein annotation entry are organized across different categories, such as function, interaction and expression, based on the type of data they contain. To provide a systematic way of categorizing computationally mapped bibliography in UniProt, we investigate a Convolution Neural Network (CNN) model to classify publications with accession annotations according to UniProtKB categories. The main challenge to categorize publications at the accession annotation level is that the same publication can be annotated with multiple proteins, and thus be associated to different category sets according to the evidence provided for the protein. We propose a model that divides the document into parts containing and not containing evidence for the protein annotation. Then, we use these parts to create different feature sets for each accession and feed them to separate layers of the network. The CNN model achieved a F1-score of 0.72, outperforming baseline models based on logistic regression and support vector machine by up to 22 and 18 percentage points, respectively. We believe that such approach could be used to systematically categorize the computationally mapped bibliography in UniProtKB, which represents a significant set of the publications, and help curators to decide whether a publication is relevant for further curation for a protein accession.
- University of Geneva Switzerland
- SIB Swiss Institute of Bioinformatics Switzerland
- University of Applied Sciences and Arts Northwestern Switzerland Switzerland
- University of Applied Sciences and Arts Western Switzerland Switzerland
- University Ucinf Chile
Proteins / genetics, Deep Learning, Knowledge Bases, Proteins, Original Article, Molecular Sequence Annotation, Databases, Protein
Proteins / genetics, Deep Learning, Knowledge Bases, Proteins, Original Article, Molecular Sequence Annotation, Databases, Protein
20 Research products, page 1 of 2
- 2017IsRelatedTo
- 2017IsRelatedTo
- 2017IsRelatedTo
- 2017IsRelatedTo
- 2017IsRelatedTo
- 2017IsRelatedTo
- 2017IsRelatedTo
chevron_left - 1
- 2
chevron_right
citations This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).8 popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.Top 10% influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).Average impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.Average
