Source-linked AI summary

AAVGen: Precision Engineering of Adeno-associated Viral Capsids for Renal Selective Targeting

Mohammadreza Ghaffarzadeh-Esfahani, Yousof Gheisari

arXiv:2602.18915v1q-bio.QMcs.AIcs.CLcs.LG

TL;DR

Kidney-directed AAV engineering is limited by restricted tropism and the challenge of optimizing multiple capsid properties simultaneously. AAVGen combines protein-language-model fine-tuning with reinforcement learning to generate capsids, producing variants with strong predicted production fitness, kidney tropism, and thermostability while preserving structural similarity.

  • Problem

    Existing AAV serotypes show limited kidney tropism and variable transduction across renal cell types, motivating precision-engineered vectors for renal targeting.

  • Method

    AAVGen combines protein-language-model fine-tuning with GSPO reinforcement learning guided by ESM-2 predictors of production fitness, kidney tropism, and thermostability.

  • Results

    99.7% of generated sequences were classified as “Best” for production fitness, 98.27% as “Good” for kidney tropism, and 88.57% as “Good” for thermostability.

  • Takeaways & Limitations

    AAVGen generated diverse capsid sequences whose predicted properties jointly improved while their structural scaffold remained similar to the baseline.

  • Takeaways & Limitations

    The generated sequences lacked in vitro validation, while limited kidney-tropism and thermostability data reduced predictive power and increased validation-set MAE.

Abstract

from arXiv · show

Adeno-associated viruses (AAVs) are promising vectors for gene therapy, but their native serotypes face limitations in tissue tropism, immune evasion, and production efficiency. Engineering capsids to overcome these hurdles is challenging due to the vast sequence space and the difficulty of simultaneously optimizing multiple functional properties. The complexity also adds when it comes to the kidney, which presents unique anatomical barriers and cellular targets that require precise and efficient vector engineering. Here, we present AAVGen, a generative artificial intelligence framework for de novo design of AAV capsids with enhanced multi-trait profiles. AAVGen integrates a protein language model (PLM) with supervised fine-tuning (SFT) and a reinforcement learning technique termed Group Sequence Policy Optimization (GSPO). The model is guided by a composite reward signal derived from three ESM-2-based regression predictors, each trained to predict a key property: production fitness, kidney tropism, and thermostability. Our results demonstrate that AAVGen produces a diverse library of novel VP1 protein sequences. In silico validations revealed that the majority of the generated variants have superior performance across all three employed indices, indicating successful multi-objective optimization. Furthermore, structural analysis via AlphaFold3 confirms that the generated sequences preserve the canonical capsid folding despite sequence diversification. AAVGen establishes a foundation for data-driven viral vector engineering, accelerating the development of next-generation AAV vectors with tailored functional characteristics.

1. Introduction

AAVs are widely used gene-delivery vectors, but kidney targeting is difficult because renal barriers and heterogeneous cell populations impede efficient transduction. AAVGen addresses this challenge by combining supervised fine-tuning, GSPO reinforcement learning, and assay-derived rewards to design kidney-tropic AAV capsids.

  • Motivation: AAVs are widely adopted clinical vectors because they are non-pathogenic, support long-term gene expression, and infect diverse cell types.Engineered capsids can also evade pre-existing neutralizing antibodies and penetrate physiological barriers.
  • Problem: Kidney targeting is challenging because the glomerular filtration barrier and heterogeneous renal cell populations hinder efficient AAV transduction.The kidney’s role in metabolic homeostasis and genetic disorders further motivates precise delivery.
  • Related work: Existing capsid-engineering strategies include recombining natural AAV variants and applying rational sequence modifications to introduce advantageous traits.Natural isolates such as AAV2, AAV8, and AAV9 have been shuffled or combined, including in Anc80L65 development.
  • Contribution: AAVGen combines a protein language model with supervised fine-tuning and Group Sequence Policy Optimization to design functionally optimized capsids for kidney tropism.Its data-driven reward functions derive from experimentally validated assays and target complex sequence–function relationships.

2. Results

AAVGen combined supervised fine-tuning and GSPO with ESM-2-based reward models to generate diverse capsid sequences optimized for production fitness, kidney tropism, and thermostability. Across 500,000 generated sequences, most variants scored highly on these objectives while retaining biologically realistic lengths and capsid-like structural similarity.

  • Model development: Three ESM-2-based regression models predicted production fitness, kidney tropism, and thermostability from AAV2 capsid sequences for use as reward functions.Production fitness reflects packaging efficiency, kidney tropism measures kidney transduction, and thermostability measures resistance to the
  • Model development: GSPO fine-tuned a ProtGPT2-based model with primary rewards for production fitness, kidney tropism, and thermostability, plus auxiliary rewards and a WT-relative reward logic mapper.Training continued until the composite reward plateaued, indicating convergence.
  • Basic analysis of generated sequences: Approximately 4% of 500,000 generated sequences were repetitive, while only 230 matched AAV2 and none matched the AAV9 WT.After duplicate removal, 1,787 sequences exactly matched sequences in the training set, indicating low exact-template replication.
  • Basic analysis of generated sequences: Generated sequences had a median length of 741 amino acids, with median sequence similarity of 99.32% and identity of 99.18%.The generated length distribution closely matched the training set, with length IQR 740–743 versus 737–743.
  • Structural analysis: AAVGen structures showed RMSD modes near 0.42 Å and 0.47 Å, compared with 0.48 Å for randomly generated baseline sequences.RMSD-length correlations were weaker for AAVGen low- and high-RMSD subgroups (Spearman ρ = 0.17 and 0.20) than for the random baseline (ρ = 0.61).

C. Joint distribution of predicted scores

Figure 5 characterizes generated AAVGen sequences through qualitative categories, pairwise correlations, and a joint three-dimensional distribution of predicted production fitness, kidney tropism, and thermostability scores. The analysis also reports that the random baseline yielded substantially lower median values across the same three metrics.

  • Joint distribution of predicted scores: Figure 5 classifies generated sequences as “Best”, “Good”, “Uncertain”, or “Bad” using predicted production fitness, kidney tropism, and thermostability scores.The categories are based jointly on the three predicted functional properties.
  • Joint distribution of predicted scores: Pairwise correlation analyses examine relationships among predicted production fitness, kidney tropism, and thermostability scores.The figure evaluates score relationships in addition to individual property classifications.
  • Joint distribution of predicted scores: A joint three-dimensional distribution visualizes predicted production fitness, kidney tropism, and thermostability scores simultaneously.The three-dimensional analysis represents the combined predicted-score space of generated sequences.
  • Joint distribution of predicted scores: −4.65, −4.19, and −4.37 were the random baseline’s median values for the same three metrics.The passage describes this as a consistent and substantial gap relative to AAVGen, supporting the functional plausibility of the generated variants.

3. Discussion

AAVGen combines supervised fine-tuning of ProtGPT2 with GSPO and multi-property regression to generate diverse, novel AAV capsids with improved multi-trait performance and preserved structural fidelity. However, limited experimental validation and scarce training data constrain definitive functional conclusions.

  • AAVGen generated a vast, diverse library of novel AAV2 VP1 variants that retained strong structural similarity to WT AAV2.
  • The framework used variable-length VP1 regression models to support generalizable, versatile AAV sequence generation.
  • Training on AAV2 and AAV9 sequences leveraged shared functional and structural features to improve model generalizability and effectiveness.
  • AAVGen learned sequence–function relationships and generated variants with superior, biologically coherent multi-trait characteristics through multi-property reward-guided reinforcement learning.
  • Generated variants showed a bimodal RMSD distribution relative to WT, suggesting that optimization explored multiple structural solutions rather than one structural optimum.
  • The sequences lacked in vitro experimental validation, while scarce high-quality kidney-tropism and thermostability data limited predictive power and increased validation-set MAE.

4. Methods · 4.1. Data collection

AAVGen was developed by integrating AAV capsid fitness data from three independent studies. The AAV2 dataset included large-scale measurements of production fitness, kidney tropism, and thermostability across distinct VP1 sequence libraries.

  • 4.1. Data collection: AAVGen integrated AAV capsid fitness data from three independent studies.
  • 4.1. Data collection: 31,579 VP1 sequences were assessed for production fitness in Ogden et al.’s AAV2 deep mutational scanning landscape.
  • 4.1. Data collection: 24,984 VP1 sequences were assessed for kidney tropism in the AAV2 fitness landscape.
  • 4.1. Data collection: 30,889 VP1 sequences were evaluated for thermostability in the AAV2 fitness landscape.
  • 4.1. Data collection: Bryant et al.’s study was incorporated to expand the mutational scope.

4.2. Data pre-processing

AAVGen’s preprocessing pipeline harmonized mutation formats, assays, and sequence representations across independent studies into model-ready datasets. Variants were reconstructed, normalized, quality-filtered, stratified into training and validation subsets, and assembled into a multi-serotype VP1 corpus for generative modeling.

  • Unified dataset construction: The pipeline unified three independent studies by reconstructing full-length VP1 sequences, normalizing fitness scores, and removing low-quality or ambiguous variants.This addressed differences in data formats, mutation types, and experimental assays.
  • Functional-property datasets: Production-fitness data combined Bryant et al. and Ogden et al. datasets, including log2-transformed scores normalized to the WT score and packaging-derived measurements.Bryant et al. covered VP1 positions 561–588, while Ogden et al. contributed deep mutational scanning read counts.
  • Functional-property datasets: Kidney-tropism processing excluded poorly covered technical replicates and filtered mouse variant counts below 10 reads or above 31,000 reads before calculating selection values.Selection values compared variant frequencies in mouse tissue samples with corresponding packaged-virus frequencies.
  • Functional-property datasets: Thermostability data were processed from packaging datasets by excluding low-coverage replicates, separately handling insertion, substitution, and deletion libraries, and removing missing or non-finite scores.The assay incubated libraries at different temperatures before digesting exposed genomes.
  • Dataset partitioning: Each dataset was stratified into training and validation subsets after grouping targets into 10 quantiles to preserve their distributions.This preparation supported training the regression models.
  • Generative-model corpus: The generative-model corpus contained 192,199 non-redundant VP1 sequences from AAV2 and AAV9, formatted in 60-character blocks with EOS tokens and split 80% for training and 20% for validation.The multi-serotype collection was designed to expose ProtGPT2 to broader viable amino acid combinations and residue-residue relationships.

4.3. AAVGen development

AAVGen combines ESM-2-based trait predictors, ProtGPT2 fine-tuning, and GSPO reinforcement learning to design AAV capsids with optimized production fitness, kidney tropism, and thermostability. Its reward system balances these objectives with sequence-length control and intra-batch uniqueness to support functional optimization and diversity.

  • Trait-prediction models: Three ESM-2 regression models predicted production fitness, kidney tropism, and thermostability to quantify functional traits for AAVGen’s multi-objective reward system.Each model was fine-tuned from an ESM-2 PLM with 8 million parameters.
  • Generative-model development: ProtGPT2 was fine-tuned on curated high-fitness AAV capsid sequences to capture residue-residue relationships across AAV serotypes.ProtGPT2 is a 36-layer decoder-only transformer with dimensionality 1,280 and approximately 738 million parameters.
  • Reinforcement learning: GSPO sampled groups of candidate sequences, increased probabilities of high-reward generations, and penalized deviations from the reference policy to stabilize optimization.The approach uses sequence-level rewards and policy updates based on group-wise candidate evaluation.
  • Reward design: The composite reward linearly combined five equally weighted functions covering production fitness, kidney tropism, thermostability, sequence-length deviation, and intra-batch uniqueness.Together, these functions promote functional optimization while maintaining generative diversity.
  • Training implementation: For each training step, AAVGen sampled G = 32 sequences with temperature = 1.0, top-p = 1.0, no top-k filtering, and a maximum completion length of 754 tokens.A repetition penalty of 1.0 was used to encourage diversity.

4.4. AAVGen Evaluation

AAVGen was evaluated at scale through generation of 500,000 protein sequences, followed by sequence-diversity, WT-similarity, functional-prediction, uncertainty-based stratification, and structural analyses. Generated variants were assessed against AAV2 WT for production fitness, kidney tropism, and thermostability, with representative variants modeled using AlphaFold3.

  • Generation pipeline: 500,000 protein sequences were generated from the fixed token “M” using sampling-based decoding with temperature = 1.0 and top-p = 1.0.Inference used batch size 64, top-k = None, and a maximum sequence length of 500 tokens.
  • Sequence-level evaluation: Sequence diversity was assessed by cumulatively sampling non-overlapping subsets of 1,000 sequences until all 500,000 generated sequences were evaluated.The protocol quantified sequence repetition among generated proteins.
  • Sequence-level evaluation: Generated proteins were globally aligned to the AAV2 WT sequence to calculate percentage identity and other sequence-based similarity metrics.Alignments used Biopython PairwiseAligner version 1.85 with match score = 2, mismatch score = −1, gap opening penalty = −2, and gap extension penalty = −0.5.
  • Functional evaluation: Regression models predicted production fitness, kidney tropism, and thermostability for all generated variants, and scores were compared with corresponding AAV2 WT values.Spearman correlation quantified relationships among predicted functional properties to assess trade-offs or co-optimization.
  • Functional evaluation: Variants were classified as Best, Good, Uncertain, or Bad according to predicted scores relative to WT and regression-model uncertainty estimated by validation-set MAE.Best variants exceeded WT by more than four times the MAE; Good variants were one to four MAE above WT; Uncertain variants were between WT and one MAE above it.
  • Structural evaluation: Representative variants spanning predicted score levels were selected for downstream structural analysis, with AlphaFold3 models compared against the AAV2 WT VP3 capsid-surface structure.Structural modeling integrated functional predictions to examine whether predicted property changes corresponded to measurable capsid-structure changes.

4.5. Hardware and training time

Training used a dedicated server with an NVIDIA V100 GPU and AMD Epyc 7502 CPU, while model-training phases required hours-scale runtimes. The reported durations ranged from 3 hours and 24 minutes for kidney regression to 11 hours and 25 minutes for fitness regression.

  • Training used a dedicated server featuring an NVIDIA V100 GPU with 32 GB of VRAM and an AMD Epyc 7502 CPU with 32 GB of RAM.
  • 11 hours and 25 minutes was required for the fitness regression model, compared with 3 hours and 24 minutes for the kidney regression model and 3 hours and 29 minutes for the thermostability regression model.
  • 9 hours and 5 minutes was required for the SFT phase of training model.

Funding

The study and its publication received no funding.

  • No funding was received for the study or its publication.

Authors contribution

M.G. led dataset preparation, model development, and model assessment, while M.G. and Y.G. jointly handled the study’s conceptualization, interpretation, manuscript preparation, and revision. All authors approved the final manuscript and accepted responsibility for the study’s integrity.

  • M.G. prepared the dataset, developed the model, and assessed the model.
  • M.G. and Y.G. jointly contributed to conceptualization, data interpretation, and drafting and revising the manuscript.
  • All authors read and approved the final version for publication and agreed to responsibility for the study’s integrity.
Loading 2602.18915v1…