Source-linked AI summary
Engaging the scientific community in high-quality biocuration: a report on the International Society for Biocuration workshop, 'Maximizing community curation for the benefit of all'
Daniela Raciti, Susan L. M. Coort, Christian Grove, Jade Hotchkiss, Matt Jeffryes, Nancy T. Li, Zhiyong Lu, Bastien Molcrette, Sushma Naithani, Maria Victoria Nugnes, Jolene Ramsey, Rene Ranzinger, Leonore Reiser, Karen E. Ross, Garrett Stevens, Courtney Thaxton, Sabrina Toro, Valerie Wood, Karen Yook, Kimberly Van Auken
TL;DR
Growing literature and declining support strain professional biocuration, motivating knowledgebases to explore community-based approaches. The paper reports on a workshop that compared practices across 18 resources and discussed common tools and best practices. It identifies community curation as an important strategy for sustaining timely biological resources, while noting continued dependence on professional engineering, curation, training, and quality assurance.
Problem
Expanding biological literature and declining knowledgebase support challenge professional curation, creating a need for additional ways to maintain current content.
Method
The workshop compared community curation pipelines across diverse resources and discussed recruitment, data types, quality assurance, incentives, machine assistance, infrastructure, and shared practices.
Results
The workshop produced a comprehensive assessment of community curation and identified shared infrastructure, tools, and proactive engagement as areas for broader development.
Takeaways & Limitations
Community curation is described as an essential strategy for sustaining biological knowledgebases as research output accelerates.
Takeaways & Limitations
High-quality community curation still depends on software engineers and professional biocurators for tools, training, and quality assurance.
Abstract
from arXiv · showhide
Biological knowledgebases traditionally rely on expert, professional curation of the research literature to maintain up-to-date collections of data organized in machine-readable form. However, despite the increasing amount of curatable biomedical knowledge, support for knowledgebases is declining, leaving these resources no alternative but to explore additional ways of updating and maintaining content. One way in which knowledgebases have addressed this problem is by engaging researchers to help curate their published papers, a process generally known as 'community curation'. As helpful as community curation can be, though, it is not universally adopted and, for groups that do have it, there is a wide range of approaches. To learn about existing community curation pipelines and explore possibilities for working towards a common approach, we organized a workshop, Maximizing Community Curation for the Benefit of All, at the 18th International Biocuration Conference, hosted by the Stowers Institute for Medical Research. Our aim was to examine the different strategies that groups use, share successes, failures, and ongoing challenges, and produce suggested deliverables for broader adoption of common best practices and tools for effective community curation. Representatives from 18 different resources, ranging from model organism and specialty knowledgebases to journals and literature resources, presented their work. The result was a comprehensive assessment of the state-of-the-art for community curation and an in-depth discussion on how community curation can become standard practice for maintaining timely, highquality biological resources that will continue to provide scientists with the essential information they need for their research.
1. Introduction
Biological knowledgebases provide structured, expertly curated data for biomedical and agricultural research, but expanding literature and declining resources strain professional biocuration. Community curation is presented as a complementary way to help knowledgebases stay current, while the workshop examined how to make it more sustainable and effective.
- Knowledgebases provide accessible, expertly curated, structured data supporting discovery, computational analysis, and clinical or other decision-making.
- Rapidly growing and increasingly complex literature has outpaced the capacity of manual professional biocuration amid limited staff and declining resources.
- Community curation engages distributed scientific contributors in identifying, extracting, annotating, and quality-controlling biological information.
- The workshop brought together knowledgebase stakeholders to discuss recruiting, training, sustaining contributors, participation barriers, and collaborative recommendations.
2. Workshop structure and goals: drawing on the collective expertise of diverse biocuration resources
The workshop drew on diverse biocuration resources through structured presentations and collaborative discussion. This process identified shared themes and created a framework for improving community curation approaches.
- The 2025 workshop included experts from 18 biocuration resources, spanning established pipelines and emerging community-based models.
- Speakers described their communities, tools, curated data types, challenges, and participation incentives, addressing both successful strategies and setbacks.
- The discussions revealed unifying themes across diverse resources and provided a framework for continued development of community curation approaches.
2.1 Participating biocuration resources
The workshop represented a broad cross-section of biocuration resources, including organism, protein, clinical, pathway, educational, genome-annotation, publishing, and literature platforms. This diversity supported exchange on scaling accessible and sustainable community engagement.
- Model organism knowledgebases: Model organism knowledgebases represented included PomBase, TAIR, and WormBase, which maintain long-standing relationships with user communities.
- Specialty knowledgebases: Other participating knowledgebases covered protein and glycoinformatics data, including UniProt, DisProt, and GlyGen.
- Clinical knowledgebases: ClinGen, SCDO, and Monarch Initiative represented knowledgebases focused on clinically relevant genetic, disease, phenotypic, and cross-species data.
- Pathway knowledgebases: Reactome, Plant Reactome, and WikiPathways represented pathway resources with expert-authored or community-driven biological pathway models.
- Education and genome annotation: CACAO, WormBase, and Plant Reactome illustrated undergraduate training and motivation approaches, while Apollo represented community-driven genome annotation.
- Publishing models: GigaScience, microPublication Biology, Genetics, and G3 illustrated publication models integrating community curation before or during publication.
- Literature resources: PubMed, PMC, Europe PMC, and BioNLP resources contributed literature repositories and tools supporting entity recognition, concept recognition, and document triage.
2.2 Key discussion topics
The workshop examined community curation across its lifecycle, from contributor recruitment and data selection to quality assurance, incentives, machine assistance, and sustainability. It aimed to assess current practice and identify improvements, collaboration opportunities, and common infrastructure.
- The workshop explored six aspects of community curation spanning contributor recruitment, curated data types, quality assurance, incentives, machine assistance, and sustainability.
- Finding and engaging community curators: Participants examined how resources recruit contributors and which engagement strategies succeed or fail.
- Data types and quality assurance: They considered which data types communities can curate to existing standards and how quality controls balance data integrity with participation barriers.
- Incentivization and recognition: The workshop addressed outreach, incentivization, and recognition models for retaining engaged community curators.
- Machine learning and artificial intelligence: It explored NLP, NER, document classification, and LLMs as ways to shift contributors toward validating machine-assisted pipelines.
- Sustainability: Participants discussed stable infrastructure, cross-resource coordination, common best practices, and a comprehensive assessment aimed at sustainable community curation.
Section 3 - Key Discussions and Takeaways
The workshop identified diverse strategies for recruiting and supporting community curators, while emphasizing persistent participation barriers, usable tools, quality assurance, and meaningful recognition. Together, these practices can help resources broaden and sustain community contributions.
- Finding and engaging community curators: Community-curation programs recruit contributors through expert networks, academic courses, research communities, and publication workflows.Examples include specialist outreach, undergraduate coursework and internships, and author-assisted curation required by several journals.
- Finding and engaging community curators: Participation remains difficult because potential contributors may lack awareness of opportunities or find curation burdensome despite visible tools and training.The barrier is especially relevant when contributors must engage with complex biological data or curation processes.
- Types of data community members curate and curation tools: User-friendly interfaces such as guided workflows and AI-assisted tools lower participation barriers by shielding contributors from complex underlying database models.The workshop connected tool usability with the ability to involve community members in curation across data types of varying complexity.
- Quality assurance: Quality assurance is essential for integrating community-contributed data reliably while preserving scientific rigor, reproducibility, accessibility, and trust.Workshop participants reported devoting substantial effort to quality control while balancing rigor against low participation barriers.
- Incentivization and recognition: Resources use formal and informal incentives—including contribution visibility, expedited processing, authorship, coursework credit, and other recognition—to encourage participation.PomBase reported that 88.5% of surveyed participants said curation increased the visibility of their work.
- Incentivization and recognition: Community curation also creates direct dialogue with knowledgebase curators and can improve data representation while developing transferable skills in data science literacy.These interactions may clarify ontology terms, identify data needing curation, and increase familiarity with metadata, formatting, and curation standards.
Section 4 - Actionable outcomes and concluding remarks
The workshop identifies shared infrastructure, coordinated prioritization, sustained professional engagement, and reliable AI integration as ways to scale community curation while preserving data quality. It also emphasizes that community curation remains dependent on professional biocurator stewardship and broader scientific outreach.
- Motivation: Declining funding and expanding biological data make scalable community curation increasingly necessary for sustaining knowledgebases.The workshop describes community curation as an essential strategy for maintaining coverage as research output accelerates.
- Shared infrastructure: Shared infrastructure could reduce duplicated development and help resources collectively identify stale, under-annotated, and high-priority data.Suggested tools include centralized dashboards, analytics, priority recommendations, and integrations with PubMed or Europe PMC.
- Engagement: 40 homepage submissions versus 446 author submissions over 15 years suggests managed, proactive outreach can outperform passive contribution links.The WormBase example is described as an approximately ten-fold increase through direct author engagement.
- Training and participation: Routine contact and training could make community curation part of the data lifecycle and sustain participation across scientists’ careers.The workshop connects early training, undergraduate involvement, transferable skills, and continued participation after career transitions or retirement.
- Automation and AI: AI is presented as essential for efficient workflows, shifting community curators toward reviewing AI-suggested annotations rather than creating annotations de novo.Effective integration requires transparent provenance, quality assurance, and mechanisms for human validation and feedback.
- Conclusions: Common priorities include centralized infrastructure, coordination across resources, collective targeting of high-value data, reliable AI pipelines, and stronger outreach through professional societies.The ISB’s Curate Now page is described as mainly reaching biocurators, motivating closer relationships with scientific societies and hands-on workshops.
- Conclusions: Community curation depends on professional biocurators to engage contributors and maintain high standards for curated data.The authors hope the workshop will support collaboration that makes high-quality community curation integral to the scientific data life cycle.
5. Author contributions
The listed authors contributed through conceptualization and writing, including drafting, review, and editing.
- Daniela Raciti is credited with conceptualization, original-draft writing, and writing review and editing.
- Susan L.M. Coort, Christian Grove, Jade Hotchkiss, Matt Jeffryes, Nancy T. Li, Zhiyong Lu, Bastien Molcrette, Sushma Naithani, and Maria Victoria Nugnes are credited with writing review and editing.
6. Group contributors list
The group contributors list identifies individuals associated with the projects and resources participating in the work.
- Contributors are listed for ACKnowledge, Apollo, CACAO, ClinGen, DisProt, NLM, Plant Reactome, Gramene, PomBase, Reactome, and the Sickle Cell Disease Ontology.
- The listed contributors include Valerio Arnaboldi, Daniela Raciti, Paul W. Sternberg, Kimberly Van Auken, Garrett Stevens, Deborah A. Siegele, Jolene Ramsey, Curtis Ross, Courtney Thaxton, Pepper St. Clair, and others.
7. Funding
The participating resources report support from governmental, philanthropic, academic, international, and institutional funders.
- DisProt, LitSuggest, LitVar, and PubTator report support through European Union projects or the NIH National Library of Medicine Intramural Research Program.
- Governmental support includes grants from NIH institutes and offices, the U.S. National Science Foundation, and UK research councils.
- European and international support includes the Wellcome Trust, European Molecular Biology Laboratory-European Bioinformatics Institute, European Union programs, ELIXIR Europe, and the European Research Council.
- Additional institutional support is reported from Texas A&M University, Maastricht University, York University, and 36 life-science research funders supporting Europe PMC.
Group
The workshop categorizes participating resources under “Model Organism Knowledgebases.”
- “Model Organism Knowledgebases” is listed as a workshop resource group.
- The group label identifies a category of participating biocuration resources.
- This category is presented as a standalone group heading.