Source-linked AI summary
Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, Huajun Chen
TL;DR
LLMs have limited proficiency in biomolecular studies, where dedicated instruction resources are lacking. Mol-Instructions provides molecule-, protein-, and text-oriented instructions and evaluates instruction-tuned models, with reported improvements in biomolecular understanding alongside recognized evaluation and specialized-generation limitations.
Problem
Biomolecular studies lack a dedicated instruction dataset for complex, interdisciplinary, and variably represented data.
Method
Mol-Instructions combines molecule-, protein-, and biomolecular-text instructions and uses them for instruction tuning across biomolecular tasks.
Results
Mol-Instructions enhances LLM molecular and protein understanding, with improvements reported across molecular metrics compared with baseline models.
Takeaways & Limitations
The dataset establishes a publicly available resource intended for continued expansion across biomolecular tasks, entries, and modalities.
Takeaways & Limitations
Common evaluation metrics capture only part of life-science output quality, while molecular generation remains below specialized smaller models.
Abstract
from arXiv · showhide
Large Language Models (LLMs), with their remarkable task-handling capabilities and innovative outputs, have catalyzed significant advancements across a spectrum of fields. However, their proficiency within specialized domains such as biomolecular studies remains limited. To address this challenge, we introduce Mol-Instructions, a comprehensive instruction dataset designed for the biomolecular domain. Mol-Instructions encompasses three key components: molecule-oriented instructions, protein-oriented instructions, and biomolecular text instructions. Each component aims to improve the understanding and prediction capabilities of LLMs concerning biomolecular features and behaviors. Through extensive instruction tuning experiments on LLMs, we demonstrate the effectiveness of Mol-Instructions in enhancing large models' performance in the intricate realm of biomolecular studies, thus fostering progress in the biomolecular research community. Mol-Instructions is publicly available for ongoing research and will undergo regular updates to enhance its applicability.
1 INTRODUCTION
Mol-Instructions addresses the lack of a dedicated biomolecular instruction dataset, a gap caused by costly data curation, broad interdisciplinary knowledge, and inconsistent biomolecular representations. It organizes molecule-, protein-, and text-oriented instructions to improve LLM understanding and prediction in biomolecular studies.
- Biomolecular applications of LLMs are promising across structural biology, computational chemistry, and drug development.
- A dedicated biomolecular dataset is needed because data curation is costly, tasks span specialized fields, and bioinformatics lacks a standardized representation.
- Mol-Instructions contains molecule-oriented, protein-oriented, and biomolecular text instructions covering chemical properties, protein functions, and biomedical information extraction.
- The dataset is collected from licensed biomolecular sources and transformed into task-specific instruction formats for domain-focused model training.
2 RELATED WORK
Prior instruction datasets relied on human annotation or automated construction, while biomolecular resources often provided limited or undisclosed instruction coverage. Mol-Instructions extends this landscape with a dedicated large-scale biomolecular dataset.
- Human-annotated instruction data offers high quality but is limited in volume, diversity, and innovation, motivating semi-automated and automated construction.
- Existing biomolecular resources include specialized datasets and models, but some instruction-data details remain undisclosed.
3 MOL-INSTRUCTIONS CONSTRUCTION
Mol-Instructions is built as a large, diverse, quality-controlled instruction resource assembled from existing biomolecular data, human-AI task creation, templates, and validation procedures. Its construction also targets text-driven protein design and reliable molecular representations.
- 3.1 UNDERLYING PRINCIPLES: Over 2 million biomolecular instructions provide broad coverage of biomolecular sequences and structures.
- 3.1 UNDERLYING PRINCIPLES: The dataset spans 17 subtasks across three biomolecule types and describes more than 11 biomolecular properties.
- 3.1 UNDERLYING PRINCIPLES: Data construction combines human-AI task-description creation, information derivation, template-based textual conversion, and quality control.
- 3.2 HUMAN-AI COLLABORATION TASK DESCRIPTION CREATION: Human-written task descriptions are expanded by GPT-3.5-turbo into varied question formulations and manually reviewed for quality.
- 3.3 INFORMATION DERIVATION FROM EXISTING DATA: Existing datasets contribute labels, inputs, outcomes, questions, and answers through direct mapping into instruction entries, while unlabeled sources are mined and augmented with AI-generated information.
- 3.4 TEMPLATE-BASED CONVERSION OF BIOLOGICAL DATA INTO TEXTUAL FORMAT: Templates convert curated protein annotations into textual design criteria, with randomly selected condition subsets accommodating input-length and training-efficiency constraints.
- 3.5 QUALITY CONTROL: Quality control removes chemically invalid SMILES strings and filters redundant protein sequences sourced mainly from curated UniProtKB/Swiss-Prot data.
4 A CLOSER LOOK AT MOL-INSTRUCTIONS
Mol-Instructions is organized around molecule, protein, and biomolecular-text tasks, with broad coverage of molecular traits, protein annotations, and descriptive text. Its analyses emphasize task diversity, biomolecular complexity, and text-based design specifications.
- 4.1 CATEGORIZATION AND POTENTIAL APPLICATIONS OF INSTRUCTION TASKS: Mol-Instructions groups tasks into molecule-oriented, protein-oriented, and biomolecular-text domains.
- 4.1 CATEGORIZATION AND POTENTIAL APPLICATIONS OF INSTRUCTION TASKS: Molecule-oriented tasks target chemical-property prediction, molecule design, and chemical-reaction accuracy across six tasks.
- 4.1 CATEGORIZATION AND POTENTIAL APPLICATIONS OF INSTRUCTION TASKS: Protein-oriented tasks cover protein domains, functions, activities, and text-directed protein design across five categories.
- 4.1 CATEGORIZATION AND POTENTIAL APPLICATIONS OF INSTRUCTION TASKS: Biomolecular-text tasks contain six information-extraction and question-answering tasks for bioinformatics and chemoinformatics literature.
- 4.2 DIVERSITY AND COMPLEXITY OF BIOMOLECULAR TRAITS: Molecular analyses examine complexity, weight, atom count, and ring count, while protein analyses cover sequence length, taxonomy, domains, gene ontology, and catalytic activity.
- 4.3 EXTENSIVE COVERAGE OF BIOMOLECULAR DESCRIPTIONS: Molecular text descriptions cover properties from basic chemical attributes to application contexts, providing multidimensional information about molecules.
- 4.3 EXTENSIVE COVERAGE OF BIOMOLECULAR DESCRIPTIONS: Protein annotations are converted into text-based specifications spanning folding, maturation, processing, and contributions to life processes.
5 EXPLORING THE POTENTIAL OF MOL-INSTRUCTIONS
Mol-Instructions is evaluated by instruction-tuning LLMs across molecule, protein, and biomolecular text tasks. The reported results show improved biomolecular-task performance, alongside evaluation and specialization limitations.
- Instruction tuning uses LLaMA-7B as the foundation model, with multiple general and domain-specific models as baselines and held-out test data for assessment.
- 5.1 INSIGHTS FROM PERFORMANCE ANALYSIS: Figure 5 compares molecule description generation with protein function, functional description, catalytic activity, and domain/motif prediction.
- 5.1 INSIGHTS FROM PERFORMANCE ANALYSIS: Life-science output evaluation remains incomplete because standard metrics capture only part of accuracy, while molecule generation still trails specialized smaller models.
- 5.1 INSIGHTS FROM PERFORMANCE ANALYSIS: Mol-Instructions improves molecular-understanding performance on every reported metric versus baseline models, including domain-specific smaller models.
- 5.1 INSIGHTS FROM PERFORMANCE ANALYSIS: LLMs generate valid molecules with greater similarity to reference molecules than baseline outputs, while also supporting molecular generation, reaction prediction, and instruction-based synthesis.
- 5.1 INSIGHTS FROM PERFORMANCE ANALYSIS: Figure 6 evaluates chemical entity recognition and chemical-disease and chemical-protein interaction extraction as biomolecular NLP tasks.
- 5.1 INSIGHTS FROM PERFORMANCE ANALYSIS: Protein-task tuning enables analysis under specific requirements and identification of fundamental protein characteristics; one generated sequence showed 40.9% identity to a UniProtKB target.
- 5.2 HARNESSING THE POWER OF MOL-INSTRUCTIONS: The authors propose cross-modal comprehension, biomolecular design exploration, and tool learning as directions for applying Mol-Instructions.
6 CONCLUSION AND FUTURE WORK
Mol-Instructions is introduced as a biomolecular instruction dataset intended to address limited resources for specialized LLM training. Future work will expand its tasks, entries, and modalities while improving biomolecular-language representation.
- Mol-Instructions is a comprehensive dataset curated for biomolecular studies to bridge gaps in existing resources and advance specialized LLM training.
- Future updates will add task types, instruction entries, and modalities aligned with advances in chemical research and AI technology.
- Current LLMs remain less proficient in biomolecular language than human language because text and biomolecule representations differ and LoRA training is limited.
- The authors identify vocabulary expansion and biomolecular encoders as possible routes for improving biomolecular-task understanding and performance.
ETHICS STATEMENT
The study reports compliance with ethical and licensing requirements for publicly available biomolecular data. It also acknowledges potential misuse risks from combining LLMs with biomolecular knowledge.
- The biomolecular data came from publicly available datasets, with no proprietary or confidential data used.
- The authors obtained permissions and licenses for incorporated third-party content.
- Quality-control and security checks were implemented to exclude harmful or malicious content from the dataset.
- The authors warn that misuse could combine LLMs and biomolecular data to generate harmful substances, including biochemical weapons or illicit drugs.
- Users are urged to follow fairness, transparency, and responsibility standards, and harmful uses are forbidden.
A ACCESSING AND UTILIZING MOL-INSTRUCTIONS
Mol-Instructions and its associated models are presented as openly accessible resources with documented data rights and usage obligations. The authors also commit to maintenance while disclaiming responsibility for problems arising from use.
- A.1 HOSTING AND ACCESS DETAILS: The dataset and models are hosted on GitHub and Hugging Face with repository guidance for exploration, structure, and use.
- A.2 DATA SOURCES AND LICENSE: Data sources and rights are documented, and sources were reviewed to ensure licenses permit the research and subsequent usage.
- A.3 USAGE GUIDELINES AND OBLIGATIONS: The authors state that all included data comply with the CC BY 4.0 License and commit to supporting users’ understanding of related obligations.
- A.3 USAGE GUIDELINES AND OBLIGATIONS: The dataset is stated to contain no personally identifiable or privacy-sensitive information, with quality and security checks against harmful content.
- A.3 USAGE GUIDELINES AND OBLIGATIONS: Users are urged to maintain fairness, transparency, and responsibility, and harmful uses are strictly forbidden.
- A.3 USAGE GUIDELINES AND OBLIGATIONS: The authors promise regular upkeep through updates, error checks, and amendments informed by field advances and user feedback.
B TASK DEFINITION AND DATA CONSTRUCTION
Mol-Instructions constructs biomolecular instruction tasks from licensed and public data sources, covering molecule, protein, and biomolecular text applications. The dataset transforms molecular, protein, and biomedical information into task-specific input-output formats for prediction and generation.
- Molecule-oriented instructions: Molecular description generation creates detailed text about a molecule’s structure, properties, biological activity, and applications from molecular descriptors.
- Molecule-oriented instructions: PubChem supplies molecular descriptions and identifiers, which are used to retrieve representations and construct molecule-oriented instruction data.
- Molecule-oriented instructions: SELFIES-description pairs are repurposed for description-guided molecule generation, producing 331,261 instruction entries.
- Scope: The dataset’s property coverage is acknowledged as narrow and not fully representative of the molecular domain.
- Protein-oriented instructions: Protein-oriented tasks include catalytic activity prediction, protein functional description generation, and domain or motif prediction from protein sequences.
- Biomolecular text instructions: Biomolecular text tasks include chemical entity recognition, chemical-disease interaction extraction, true-or-false questions, open questions, and question-answer generation from biomedical abstracts.
D EXPERIMENTAL SETUP DETAILS
The experiments instruction-tune 7B LLaMA models on separate molecule-, protein-, and text-oriented datasets, using task-specific preprocessing, optimization, and evaluation procedures. Test sets contain approximately 1k samples per task, while the remaining data are split into training and validation sets.
- Model and data setup: Three 7B LLaMA models are fine-tuned separately for molecule-oriented, protein-oriented, and biomolecular text tasks.
- Model and data setup: Approximately 1k samples per task are reserved for testing, with remaining samples divided into training and validation sets at an 8:2 ratio.
- Preprocessing: Protein amino acid sequences and molecular SELFIES strings are tokenized as human language using the LLaMA byte-pair encoding model.
- Optimization: LoRA fine-tuning is used for molecule- and text-oriented datasets, whereas protein-oriented data use full fine-tuning with ZeRO memory optimization.
- Evaluation: Molecular understanding is evaluated with BLEU, ROUGE, and METEOR, while molecule generation additionally checks validity with RDKit and exact match.
- Evaluation: Protein functional descriptions are evaluated with ROUGE, and biomolecular text tasks use general metrics for question answering, entity recognition, and relation extraction.
F ADDITIONAL RESULTS
Additional experiments report that Mol-Instructions improves task-specific performance across molecule-oriented, protein-oriented, and biomolecular text tasks. The authors attribute the gains to instructions that align outputs with specialized task requirements, while noting that the current models remain preliminary demonstrations.
- Molecule-oriented tasks: Instruction-tuned models outperform or better align with task requirements than models lacking specialized instruction tuning across molecular tasks.
- Protein-oriented tasks: Fine-tuned models provide consistent, user-specific protein functional analysis, including catalytic activity classification examples.
- Biomolecular text tasks: Mol-Instructions helps identify chemical entities and relations and produces more detailed, structured answers for biomolecular question-answering tasks.
- Overall findings: Mol-Instructions improves LLM execution of molecule-oriented, protein-oriented, and biomolecular text tasks.
- Scope of results: The authors characterize the instruction-tuned models as preliminary demonstrations with limited applicability to real-world production tasks.
G LIMITATIONS AND ETHICAL CONCERNS
The paper identifies both technical and ethical boundaries for Mol-Instructions. Current models have limited production-level applicability, while biomolecular generation capabilities could be misused and therefore require responsible deployment practices.
- Ethical concerns: LLM capabilities for forecasting reactions and designing molecules could be misused to synthesize harmful biochemical agents or bioweapons.
- Ethical concerns: The paper recommends stringent usage guidelines and monitoring mechanisms for publicly available generative AI tools in bioengineering.
- Safeguards: Suggested safeguards include regulated access, monitoring usage patterns, community oversight, and transparent reporting.
- Safeguards: The authors emphasize caution, diligence, and ethical principles when applying LLMs in sensitive bioscience domains.