Source-linked AI summary
An expanded evaluation of protein function prediction methods shows an improvement in accuracy
Yuxiang Jiang, Tal Ronnen Oron, Wyatt T Clark, Asma R Bankapur, Daniel D'Andrea, Rosalba Lepore, Christopher S Funk, Indika Kahanda, Karin M Verspoor, Asa Ben-Hur, Emily Koo, Duncan Penfold-Brown, Dennis Shasha, Noah Youngs, Richard Bonneau, Alexandra Lin, Sayed ME Sahraeian, Pier Luigi Martelli, Giuseppe Profiti, Rita Casadio, Renzhi Cao, Zhaolong Zhong, Jianlin Cheng, Adrian Altenhoff, Nives Skunca, Christophe Dessimoz, Tunca Dogan, Kai Hakala, Suwisa Kaewphan, Farrokh Mehryary, Tapio Salakoski, Filip Ginter, Hai Fang, Ben Smithers, Matt Oates, Julian Gough, Petri Törönen, Patrik Koskinen, Liisa Holm, Ching-Tai Chen, Wen-Lian Hsu, Kevin Bryson, Domenico Cozzetto, Federico Minneci, David T Jones, Samuel Chapman, Dukka B K. C., Ishita K Khan, Daisuke Kihara, Dan Ofer, Nadav Rappoport, Amos Stern, Elena Cibrian-Uhalte, Paul Denny, Rebecca E Foulger, Reija Hieta, Duncan Legge, Ruth C Lovering, Michele Magrane, Anna N Melidoni, Prudence Mutowo-Meullenet, Klemens Pichler, Aleksandra Shypitsyna, Biao Li, Pooya Zakeri, Sarah ElShal, Léon-Charles Tranchevent, Sayoni Das, Natalie L Dawson, David Lee, Jonathan G Lees, Ian Sillitoe, Prajwal Bhat, Tamás Nepusz, Alfonso E Romero, Rajkumar Sasidharan, Haixuan Yang, Alberto Paccanaro, Jesse Gillis, Adriana E Sedeño-Cortés, Paul Pavlidis, Shou Feng, Juan M Cejuela, Tatyana Goldberg, Tobias Hamp, Lothar Richter, Asaf Salamov, Toni Gabaldon, Marina Marcet-Houben, Fran Supek, Qingtian Gong, Wei Ning, Yuanpeng Zhou, Weidong Tian, Marco Falda, Paolo Fontana, Enrico Lavezzo, Stefano Toppo, Carlo Ferrari, Manuel Giollo, Damiano Piovesan, Silvio Tosatto, Angela del Pozo, José M Fernández, Paolo Maietta, Alfonso Valencia, Michael L Tress, Alfredo Benso, Stefano Di Carlo, Gianfranco Politano, Alessandro Savino, Hafeez Ur Rehman, Matteo Re, Marco Mesiti, Giorgio Valentini, Joachim W Bargsten, Aalt DJ van Dijk, Branislava Gemovic, Sanja Glisic, Vladmir Perovic, Veljko Veljkovic, Nevena Veljkovic, Danillo C Almeida-e-Silva, Ricardo ZN Vencio, Malvika Sharan, Jörg Vogel, Lakesh Kansakar, Shanshan Zhang, Slobodan Vucetic, Zheng Wang, Michael JE Sternberg, Mark N Wass, Rachael P Huntley, Maria J Martin, Claire O'Donovan, Peter N Robinson, Yves Moreau, Anna Tramontano, Patricia C Babbitt, Steven E Brenner, Michal Linial, Christine A Orengo, Burkhard Rost, Casey S Greene, Sean D Mooney, Iddo Friedberg, Predrag Radivojac
TL;DR
Assigning function to proteins is a major bottleneck, while evaluating computational predictors and tracking field-wide progress remain challenging. CAFA2 expanded the assessment of protein-function prediction, and its results showed encouraging progress while identifying substantial room for improvement.
Problem
Assigning function to proteins is a major bottleneck because experiments are reliable but relatively low-throughput, making computational prediction increasingly important; evaluating predictors and tracking progress remain challenging.
Method
CAFA2 evaluated computational protein-function predictors through expanded datasets, ontologies, evaluation scenarios, and performance metrics, including protein-centric and term-centric assessments.
Results
CAFA2's top methods showed encouraging progress in MFO and BPO, although raw scores remained substantially improvable, particularly for BPO, CCO, and HPO.
Takeaways & Limitations
The results provide information on the state of protein-function prediction, can guide concept-annotation method development, and may help experimental studies prioritize targets.
Takeaways & Limitations
Raw scores still leave significant room for improvement across all ontologies, especially BPO, CCO, and HPO.
Abstract
from arXiv · showhide
Background: The increasing volume and variety of genotypic and phenotypic data is a major defining characteristic of modern biomedical sciences. At the same time, the limitations in technology for generating data and the inherently stochastic nature of biomolecular events have led to the discrepancy between the volume of data and the amount of knowledge gleaned from it. A major bottleneck in our ability to understand the molecular underpinnings of life is the assignment of function to biological macromolecules, especially proteins. While molecular experiments provide the most reliable annotation of proteins, their relatively low throughput and restricted purview have led to an increasing role for computational function prediction. However, accurately assessing methods for protein function prediction and tracking progress in the field remain challenging. Methodology: We have conducted the second Critical Assessment of Functional Annotation (CAFA), a timed challenge to assess computational methods that automatically assign protein function. One hundred twenty-six methods from 56 research groups were evaluated for their ability to predict biological functions using the Gene Ontology and gene-disease associations using the Human Phenotype Ontology on a set of 3,681 proteins from 18 species. CAFA2 featured significantly expanded analysis compared with CAFA1, with regards to data set size, variety, and assessment metrics. To review progress in the field, the analysis also compared the best methods participating in CAFA1 to those of CAFA2. Conclusions: The top performing methods in CAFA2 outperformed the best methods from CAFA1, demonstrating that computational function prediction is improving. This increased accuracy can be attributed to the combined effect of the growing number of experimental annotations and improved methods for function prediction.
January 6, 2016
Protein function assignment is a major bottleneck because experiments are reliable but low-throughput, motivating computational prediction and rigorous assessment. CAFA2 expanded evaluation and found improved performance over CAFA1.
- Protein function assignment remains difficult because experimental annotation is reliable but low-throughput and computational methods are challenging to assess.
- CAFA2 evaluated 126 methods from 56 groups on 3,681 proteins across 18 species using Gene Ontology and Human Phenotype Ontology tasks.
- CAFA2 expanded analysis relative to CAFA1 in dataset size, variety, and assessment metrics.
- CAFA2 methods outperformed the best CAFA1 methods, demonstrating improvement in computational protein function prediction.
Introduction
CAFA2 followed CAFA1 to measure progress in protein function prediction while broadening the ontologies, targets, evaluation scenarios, and metrics. The expanded design supported comparison with earlier state-of-the-art methods.
- CAFA2 aimed to quantify progress while expanding targets, ontologies, analysis scenarios, and evaluation metrics.
- CAFA1 showed that advanced methods integrating multiple sequence hits and data types outperformed local sequence-similarity transfer.
- CAFA1 also exposed challenges involving experimental techniques, protein systems data, and evaluation metrics.
- CAFA2 was organized three years after CAFA1 to evaluate progress in function prediction and expand assessment to new ontologies and performance analyses.
Methods
CAFA2 assessed protein-function predictors across expanded ontology benchmarks, evaluation modes, and metrics. Its methods distinguish protein-centric from term-centric prediction and use curated annotations with multiple performance measures.
- Data sets: CAFA2 evaluated predictions for Gene Ontology and Human Phenotype Ontology, adding Cellular Component and disease-term association tasks beyond CAFA1.
- Experiment overview: The study included 126 methods from 56 groups, with 121 methods submitting Gene Ontology predictions and 30 participating in Human Phenotype Ontology disease-gene tasks.
- Evaluation: Protein-centric evaluation ranks ontology terms for each protein, whereas term-centric evaluation classifies whether a specific term applies to a protein.
- Partial and full evaluation modes: Partial evaluation scores selected target subsets, while full evaluation scores all benchmark proteins and penalizes missing predictions.
- Evaluation metrics: Protein-centric assessment used precision-recall and remaining uncertainty-misinformation curves, summarized with Fmax and minimum semantic distance.
- Data sets: Benchmarks used experimental annotations from Swiss-Prot, UniProt-GOA, the Gene Ontology consortium, and the Human Phenotype Ontology database.
Results
CAFA2 performance varied across ontologies and evaluation perspectives, with molecular-function prediction strongest and phenotype prediction difficult in protein-centric evaluation. Compared with CAFA1, top CAFA2 methods improved, while method diversity differed across ontologies and ensemble gains may be greatest where approaches are less similar.
- Overall performance: Predictor accuracy differed across ontologies because ontology structure, term predictability, training-set size, and annotation biases varied.The reported factors include ontology size, depth, branching factor, abstraction, and annotation status.
- Overall performance: MFO was the strongest ontology, with top methods reaching Fmax around 0.6, whereas the best BPO method scored slightly below 0.4.BPO performance remained below MFO performance, consistent with the CAFA1 pattern.
- Overall performance: Top predictors did not exceed the Naïve method under Fmax in the newly added ontologies, but slightly surpassed it under Smin for CCO.Frequent general CCO terms can boost Naïve Fmax, while semantic-distance weighting penalizes those less informative predictions.
- Term-centric evaluation: Term-centric analysis showed that most HPO terms exceeded the Naïve model, with some terms achieving AUCs around 0.7 despite poor overall protein-centric performance.The analysis averaged AUC across terms with at least ten positive sequences and identified both specific and general phenotype terms as well predicted.
- Easy vs. difficult benchmarks: Benchmark difficulty affected BLAST strongly, but most top methods were insensitive to easy versus difficult categories, supporting integration of multiple data sources and unreliable hits.Easy and difficult categories were defined using a 60% maximum global sequence-identity cutoff.
- Top methods have improved since CAFA1: All top CAFA2 methods outperformed their CAFA1 counterparts, indicating improvement beyond baseline changes and sequence-transfer effects.Baseline performance generally did not improve, except for BLAST on BPO with newer Swiss-Prot; the authors suggest growing annotations and improved methods jointly contributed.
Conclusions
CAFA2 showed measurable progress in computational protein-function prediction, while revealing that performance depends on ontology, benchmark, and evaluation metric. The assessment also highlights persistent challenges in fair evaluation and broad improvement across all ontologies.
- CAFA2 expanded assessment across targets, ontologies, analysis scenarios, and metrics to quantify progress in function prediction.The experiment was designed to assess the field’s status and provide information useful for developing annotation methods and prioritizing experimental studies.
- All top CAFA2 methods outperformed their CAFA1 counterparts in head-to-head comparisons, indicating progress in protein-function prediction.The authors attribute the improvement mainly to new methods, while acknowledging that larger training sets and methodological novelty are difficult to disentangle.
- Different metrics can produce different interpretations, especially for cellular component and human phenotype ontology evaluations.Protein-centric and term-centric views generally correlated for molecular function and biological process but could diverge substantially for cellular component and human phenotype.
- No single method dominated every benchmark because rankings changed with ontology, benchmark type, evaluation setting, and metric.A small group of methods performed generally well, but specialization and differing application objectives affected rankings.
- Methods generally performed better in partial evaluation mode when allowed to select a reliable subset of predictions, whereas limited-knowledge evaluation provided no accuracy boost.The limited-knowledge category was new, and few methods may have been optimized for it.
- Automated annotation remains challenging, with significant room for improvement particularly in biological process, cellular component, and human phenotype categories.Future evaluations should develop experiment-driven components to address limitations of term-centric assessment.
Authors’ contributions
The project combined leadership, analysis, data acquisition, software, and biocuration contributions across the author team.
- The authors’ contributions spanned experiment supervision, analysis, data acquisition, prediction storage, web-interface development, and biocuration.
Additional Files
The paper provides supplementary analyses, complete prediction data, and the code used for CAFA2.
- Three supplementary resources provide additional analyses, a repository of data and full method predictions, and the complete CAFA2 codebase.The code is available through the cited GitHub repository.
Supplementary Information
The supplementary information documents benchmark annotation structure, predictor comparisons, species and difficulty breakdowns, and alternative evaluation metrics. It also defines weighted precision-recall calculations and provides additional data and code resources.
- Supplementary resources: The supplementary package includes participating-team and keyword tables, additional analyses and full prediction results, and code used in CAFA2.The additional data are described as 297MB.
- Top predictors: The supplementary analyses compare top predictors using precision-recall curves, Fmax bars, weighted precision-recall, and normalized remaining uncertainty-misinformation across ontologies and benchmark categories.Comparisons cover ontology type, benchmark difficulty, organism group, and species-specific performance.
- Benchmark sequence identity: Benchmark proteins are grouped as easy or difficult according to whether their most similar experimentally annotated template has at least 60% global sequence identity.The sequence-identity histograms define the grouping threshold used in the supplementary comparisons.
- Species breakdown: The supplementary figures report predictor performance by species, including only species with at least 15 benchmark proteins for the Fmax analysis.The species breakdown spans Molecular Function, Biological Process, and Cellular Component ontologies.
- Weighted evaluation: Weighted precision and recall assign each ontology term a weight based on its information content, with precision and recall computed over predicted and true term sets at threshold τ.The full evaluation uses all benchmark proteins, whereas partial evaluation uses proteins with at least one qualifying prediction.