Source-linked AI summary
Gene Ontology: Pitfalls, Biases, Remedies
Pascale Gaudet, Christophe Dessimoz
TL;DR
The chapter examines how incomplete, heterogeneous, and structurally complex GO data can mislead functional analyses. It explains the ontology’s relation semantics, annotation qualifiers, and major biases, then recommends controlling confounders and checking conclusions across updated releases.
Problem
GO annotations are heterogeneous, incomplete, and observational, while qualifiers and ontology relations can alter their interpretation and bias aggregate analyses.
Method
The chapter surveys common GO pitfalls and biases, explains their mechanisms, and presents remedies and best practices.
Results
The chapter identifies annotation incompleteness, relation semantics, qualifiers, species differences, authorship, and annotator biases as important considerations for GO analyses.
Takeaways & Limitations
Users can make better use of GO by understanding its subtleties, controlling known confounders, seeking unknown ones, and proceeding cautiously.
Takeaways & Limitations
GO is necessarily incomplete, so absence of evidence of function does not imply absence of function.
Abstract
from arXiv · showhide
The Gene Ontology (GO) is a formidable resource but there are several considerations about it that are essential to understand the data and interpret it correctly. The GO is sufficiently simple that it can be used without deep understanding of its structure or how it is developed, which is both a strength and a weakness. In this chapter, we discuss some common misinterpretations of the ontology and the annotations. A better understanding of the pitfalls and the biases in the GO should help users make the most of this very rich resource. We also review some of the misconceptions and misleading assumptions commonly made about GO, including the effect of data incompleteness, the importance of annotation qualifiers, and the transitivity or lack thereof associated with different ontology relations. We also discuss several biases that can confound aggregate analyses such as gene enrichment analyses. For each of these pitfalls and biases, we suggest remedies and best practices.
1. Introduction
GO enables large-scale functional analyses, but its heterogeneous, observational, dynamic, and incomplete annotations can bias conclusions unless users control confounders and account for missing knowledge.
- 1. Introduction: GO annotations are observational and heterogeneous, so unknown or unmeasured confounders can bias large-scale analyses.The risk is amplified when heterogeneous data are aggregated across many comparisons.
- 1.1 Simpson’s Paradox: the perils of data aggregation: Because aggregated GO data can exhibit Simpson’s paradox, apparent associations may differ from patterns within individual datasets.This makes control of potential biases and confounders important.
- 1.2 The inherent incompleteness of the Gene Ontology (Open World Assumption): GO’s incompleteness means that absence of an annotation does not imply absence of biological function.This is the Open World Assumption.
- 1.2 The inherent incompleteness of the Gene Ontology (Open World Assumption): Incomplete annotations can inflate false-positive rates when computational function-prediction methods are evaluated.Thinning annotations or comparing successive database releases can gauge the effect of incompleteness.
2. Gene Ontology structure
GO analysis depends on respecting the ontology’s graph structure, relation semantics, inferred links, and changing annotation background distributions.
- 2. Gene Ontology structure: GO’s uneven term granularity creates the shallow annotation problem and complicates semantic similarity measurements.Information-theoretic similarity measures can partly mitigate this problem.
- 2. Gene Ontology structure: Transitive “is a” and “part of” relations propagate annotations to parent terms, preserving the associated function at a more general level.A serine/threonine protein kinase activity annotation therefore implies protein kinase activity.
- 2.1 Understanding relationships between ontological concepts: Non-transitive relations such as “regulates” block valid parent-term inferences, so peptidase inhibitor activity does not imply a role in proteolysis.Logical reasoning over GO must therefore account for relation transitivity.
- 2.2 Inter-ontology links and their impact on GO enrichment analyses: Cross-aspect “part of” links can automatically infer annotations with the same evidence and reference, increasing coverage but affecting enrichment analyses.For example, DNA ligase activity can infer DNA ligation.
- 2.2 Inter-ontology links and their impact on GO enrichment analyses: Sudden changes in annotation counts alter enrichment background distributions, so analyses should use current releases and test conclusions across recent versions.The ATPase activity example illustrates strong temporal annotation variation.
3. Gene Ontology annotations
GO annotations contain qualifiers and literature-supported evidence that can materially change how gene–term associations should be interpreted. Negative and apparently contradictory annotations may reflect context-dependent biology rather than errors.
- GO qualifiers modify the meaning of gene-product–term associations, including “NOT”, “contributes to”, and “co-localizes with”.
- “Contributes to” can annotate complex subunits to a molecular function even when some subunits do not directly perform the activity.This can make a cyclin annotated with protein kinase activity appear unintuitive.
- “Co-localizes with” may indicate transient or peripheral association, or uncertainty about bona fide complex membership, and the annotation does not distinguish these meanings.
- “NOT” records evidence that a gene product lacks a function, but positive and negative annotations can coexist for the same gene and term.
- The ARR2 contradiction reflects differing experimental findings and conditions, illustrating that GO disagreements can mirror context-dependent primary literature rather than annotation mistakes.Differences may involve tissue, subcellular localization, time, or experimental growth conditions.
- A negative annotation can flag proteins whose sequences resemble active homologs even when their biological role differs, as illustrated by the STRADA pseudokinase.
3.3 Annotation extensions
Annotation extensions add contextual information to GO associations, making them more expressive than simple gene-product–term links. Their novelty and inconsistent tool support require care in downstream analyses.
- Annotation extensions add context such as target genes, activity locations, dynamic localization, and enzyme substrates to GO annotations.Examples include opsin-4 activity in retinal ganglion cells and bir1 localization during mitotic anaphase.
- Extension data are available through AmiGO, QuickGO, and GAF2.0 annotation files.
- Because extension guidelines are still developing, usage varies across databases and most tools do not yet incorporate the information.
- Extensions create virtual GO classes that can span multiple actual classes and parent lineages, potentially inflating annotation counts in enrichment analyses.
3.4 Biases associated with particular evidence codes
Evidence codes and annotation methods differ in precision, confidence, specificity, and information content. These differences can bias similarity and enrichment analyses, so evidence provenance and annotation multiplicity require normalization and careful interpretation.
- Evidence codes inform interpretation but generally cannot be used directly to exclude low-confidence data because most experiment types lack quantitative confidence measures.
- Mutant phenotype and genetic-interaction annotations may weakly identify how a gene product is implicated because phenotypic evidence is inherently derivative.
- Physical interactions support molecular-function and biological-process annotations mainly through low-confidence guilt by association, while expression-pattern inferences are typically low confidence.
- High-throughput annotations tend toward high-level terms and limited functions, reducing term information content and affecting similarity analyses.Such analyses contributed as much as 25% of GO annotations in the cited report.
- Automatic annotations use diverse methods such as domains, enzyme classifications, BLAST, and orthology, recorded through reference codes rather than evidence codes.
- Multiple annotations to one term can corroborate evidence but may artificially inflate term frequencies, while experiments with different specificity can place annotations at different levels.
3.5 Differences among species
GO annotation coverage and emphasis differ substantially among species because research priorities vary. Using an undifferentiated database background can therefore make apparent enrichments reflect species-specific annotation patterns.
- Species differ substantially in the nature and extent of GO annotations, reflecting distinct research emphases such as zebrafish development and rat toxicology.
- Using the entire database as an enrichment background can make zebrafish interaction partners appear enriched for developmental genes because annotation coverage is species-biased.
3.6 Authorship bias
GO annotations from the same paper are more similar, creating an authorship bias that can falsely suggest stronger functional conservation among same-species paralogs than orthologs.
- Controlling for authorship made the apparent functional-similarity advantage of same-species paralogs disappear and favor orthologs.Same-species paralog annotations were ~50 times more likely than ortholog annotations to come from the same paper.
- Annotations from the same article were more similar on average than annotations from different papers.
- The difference remained smaller but significant when comparing different papers with shared authors against papers with no authors in common.
3.7 Annotator bias
GO annotations vary systematically across curators and evidence types, so apparent functional differences can reflect annotator or propagation biases rather than biology alone.
- MGI and UniProt annotated different phenotype-supported GO terms, indicating that each group emphasizes specific biological aspects rather than representing species-wide function uniformly.MGI’s top IMP term was “in utero embryonic development” with 1170 annotations, whereas UniProt emphasized “regulation of circadian rhythm.”
- Electronic annotations show higher average similarity among homologs than experimental annotations, with curated annotations intermediate.Electronic annotations are typically propagated among homologous sequences, which can increase homolog functional similarity.
- Different evidence-code distributions and curator practices can bias analyses, especially when comparing model organisms with non-model organisms.Non-model-organism annotations are likely to contain more electronically propagated annotations.
- Functional-conservation analyses across species divergence can be confounded because computational methods preferentially infer function among phylogenetically close homologs.
3.9 Imbalance between positive and negative annotations
Uneven coverage of positive and negative annotations creates evaluation bias, particularly for machine-learning methods assessed under incomplete negative labels.
- Missing negative annotations can strongly alter the ranking produced by ROC-area and precision-recall-area evaluation metrics.The reported study examined different levels of missing negative annotations and found substantial effects on metric rankings.
- The imbalance between positive and negative annotations complicates machine-learning training and may have an even larger effect under the open-world assumption.
4 Getting help
Correct GO interpretation requires accounting for annotation qualifiers and evidence-code distributions, while using available consortium resources for guidance.
- Qualifiers such as “NOT” and “co-localizes with” fundamentally change annotation meaning and should be handled by analysis tools and software libraries.
- Evidence codes differ in scope, specificity, and abundance, so statistical analyses should consider their distribution and control for it when needed.
- The chapter points users toward resources provided by the GO consortium for additional help.
5 Conclusion
The chapter surveys major Gene Ontology pitfalls and biases, emphasizing that many can be addressed through careful interpretation and practical remedies.
- Understanding GO subtleties and controlling both known and suspected confounders enables users to make better use of this resource.The chapter recommends proceeding cautiously because GO data are observational and inherently subject to bias.
- Many GO pitfalls and biases have simple remedies despite the breadth of potential issues.The chapter summarizes these issues and remedies in Table 1.