Source-linked AI summary
IGoR: a tool for high-throughput immune repertoire analysis
Quentin Marcou, Thierry Mora, Aleksandra M Walczak
TL;DR
High-throughput repertoire sequencing requires accurate characterization and interpretation of receptor-generation statistics. IGoR learns these statistics from cDNA or gDNA, probabilistically evaluates alternative recombination scenarios, and models B-cell hypermutation. It captures scenario dependencies and generation probabilities, predicts complete recombination scenarios about 2.5 times better than existing methods, and identifies sample-size and hypermutation-rate boundaries for reliable inference.
Problem
High-throughput immune repertoire sequencing needs accurate methods to characterize, analyze, and interpret B- and T-cell receptor-generation datasets.
Method
IGoR learns V(D)J recombination and somatic-hypermutation statistics, then probabilistically enumerates and ranks possible scenarios for each sequence.
Results
2.5 times better: IGoR identified complete recombination scenarios more accurately than MiXCR and Partis, while capturing dependencies such as impossible TRB D2-J1 pairs.
Takeaways & Limitations
Summing probabilities across scenarios provides sequence-generation probabilities that can indicate convergent recombination in antigen-specific or autoimmune-sequence studies.
Takeaways & Limitations
Partis does not include palindromic insertions, and its performance becomes comparable to MiXCR when evaluation excludes sequences containing them.
Abstract
from arXiv · showhide
High throughput immune repertoire sequencing is promising to lead to new statistical diagnostic tools for medicine and biology. Successful implementations of these methods require a correct characterization, analysis and interpretation of these datasets. We present IGoR -- a new comprehensive tool that takes B or T-cell receptors sequence reads and quantitatively characterizes the statistics of receptor generation from both cDNA and gDNA. It probabilistically annotates sequences and its modular structure can investigate models of increasing biological complexity for different organisms. For B-cells IGoR returns the hypermutation statistics, which we use to reveal co-localization of hypermutations along the sequence. We demonstrate that IGoR outperforms existing tools in accuracy and estimate the sample sizes needed for reliable repertoire characterization.
RESULTS
IGoR learns recombination statistics, probabilistically ranks alternative sequence-generation scenarios, and generates sequences from specified statistics. Its analyses show reproducible inference, sample-size and mutation-rate boundaries, substantial scenario degeneracy, and the value of summing scenario probabilities.
- Probabilistic assignment of recombination scenarios: IGoR enumerates possible recombination and hypermutation scenarios, weights them by likelihood, and operates in learning, analysis, and generation modes.Learning uses Sparse Expectation-Maximization; analysis ranks scenarios and reports generation probability; generation produces synthetic sequences.
- Inference of V(D)J recombination: The same TRB insertion and deletion distributions were inferred across individuals, laboratories, protocols, and DNA or mRNA data.V and J gene usage instead varied across individuals and sequencing technologies, suggesting primer-dependent biases.
- Inference of V(D)J recombination: ∼5000 unique out-of-frame sequences sufficed to learn an accurate TRB model at error rate 10^-3, while hypermutation rates of 10^-1 significantly degraded accuracy.Most estimation error arose from deletion profiles; the result supports using naive, non-hypermutated B-cell sequences for BCR statistics.
- Analysis of scenario degeneracy: The maximum-likelihood scenario was incorrect in 72% of IGH sequences and 85% of 60bp TRB sequences, revealing substantial recombination degeneracy.The true scenario often appeared far down the likelihood ranking.
- Analysis of scenario degeneracy: At least 30 to 50 ranked scenarios were required to explain 95% of total sequence likelihood.The analysis therefore considers multiple scenarios to characterize the generation process rather than relying on a single assignment.
- Analysis of scenario degeneracy: IGoR computes sequence-generation probability by summing the probabilities of all possible scenarios, enabling an indicator of convergent recombination.The passage connects generation probability to sharing properties between healthy individuals and studies of antigen-specific or autoimmune-related sequences.
Comparison to other methods
IGoR was compared with MiXCR and Partis on synthetic IGH and TRB data. It better captures dependencies and insertion statistics while predicting complete recombination scenarios about 2.5 times better, although Partis excludes palindromic insertions.
- Comparison to other methods: 2.5 times better: IGoR predicted the complete recombination scenario and each individual component more accurately than MiXCR and Partis.Performance was evaluated against known generation scenarios in synthetic data.
- Comparison to other methods: Partis excludes palindromic insertions; restricting evaluation to sequences without them made Partis’ performance comparable to MiXCR’s.IGoR and MiXCR represent these insertions by appending a short palindromic sequence to germline segments.
- Comparison to other methods: IGoR, MiXCR, and Partis produced similar V and J gene-usage and deletion profiles, but only IGoR captured the TRB D-J dependency.IGoR assigned zero frequency to impossible D2-J1 pairs, whereas MiXCR assigned 20% of events to them.
- Comparison to other methods: IGoR accurately reconstructed insertion distributions, while the other methods systematically over- or under-estimated them.The comparison used statistics learned or assigned from synthetic 130-base-pair IGH sequences.
Somatic hypermutations
IGoR models somatic hypermutations with sequence-dependent PWMs and uses them to analyze mutation patterns in memory-B-cell IGH sequences. The inferred motifs are reproducible and reveal segment-specific mutation rules, co-localization, and limits of context-only prediction.
- Model: IGoR infers sequence-dependent hypermutation rates using position-weight matrices learned from memory-B-cell out-of-frame IGH sequences.The model uses immediate n-mer context and was trained with Expectation-Maximization.
- Model performance: The PWM predicts V-gene mutation rates with correlation r = 0.7, while motifs are highly reproducible across individuals at r = 0.98.Overall mutation rates nevertheless differed two-fold between the two individuals.
- Segment specificity: The inferred PWM contributions extend at least four nucleotides from the mutation site, producing detailed motifs that cannot be reduced to a few simple sequence patterns.This may reflect broad context dependence or indirect capture of non-contextual effects.
- Segment specificity: Mutation motifs differ substantially among V, D, and J genes, and cross-segment prediction falls from r = 0.7 for V genes to r = 0.24 for J-learned predictions of V-gene rates.The results indicate that context-dependent motifs alone do not explain all hypermutation variability.
- Co-localization: Observed mutation-count distributions are more skewed and longer-tailed than predictions assuming independent hypermutations under the inferred PWM.The discrepancy is consistent with variable numbers of affinity-maturation cycles among B cells.
- Co-localization: Correlated hypermutations have an estimated co-localized-region length of about 15 base pairs.Synthetic sequences generated under the model have g(r) ≈1 by construction, providing the comparison baseline.
- Sequence analysis: IGoR estimates individual-sequence generation probabilities accurately despite uncertainty about highly hypermutated ancestral sequences, achieving r = 0.97 on synthetic data.It evaluates possible recombination and hypermutation scenarios and their potential ancestral sequences.
DISCUSSION
The discussion presents IGoR as a probabilistic, modular framework that improves recombination inference and supports flexible repertoire modeling. It also highlights synthetic-sequence controls, broader hypermutation structure, and boundaries imposed by the model and training data.
- DISCUSSION: Probabilistic germline alignments correct systematic biases in V(D)J recombination statistics and improve recombination-scenario prediction over previous methods.The approach evaluates multiple possible alignments and scenarios rather than relying only on deterministic assignments.
- DISCUSSION: More than 70% of sequences have an incorrectly called recombination scenario even with a perfect estimator, cautioning against deterministic assignments.The result reflects intrinsic ambiguity in receptor generation, not merely imperfect software performance.
- DISCUSSION: IGoR’s configurable dependency structure supports TCRs and immunoglobulins across species with available genomic data, including unusual or incomplete rearrangements.Examples include D-J, DD2/DD3, and hybrid TRA/TRD rearrangements.
- DISCUSSION: Synthetic sequences generated from learned models can serve as controls for disease-association studies and help distinguish antigen-specific clonotypes from public convergent sequences.This could reduce reliance on a healthy control cohort within the stated application.
- DISCUSSION: Distinct V-, D-, and J-segment motifs and hypermutation co-localization suggest that hotspot formation reflects context, position-specific effects, and nearby mutation co-occurrence.Future prediction improvements require better understanding of AID operation.
- Recombination model: The framework assumes recombination-feature dependencies can be represented by an acyclic graph, or Bayesian network, configured through setup files.The model includes segment choices, deletions, insertions, and dependencies among these stochastic features.
Context dependent hypermutation model
IGoR models receptor errors and memory-B-cell hypermutations probabilistically while enumerating plausible recombination and mutation scenarios. It combines context-dependent mutation probabilities with scenario likelihoods to estimate sequence and generation probabilities.
- Model definition: TCR and naive-BCR processing assumes a constant sequencing-error probability, whereas memory-BCR processing uses context-dependent hypermutation rates.The context-dependent model applies along the V, D, and J genes.
- Model definition: Each mutation probability depends on the surrounding (2m + 1)-mer context through additive position-weight-matrix contributions and an overall mutation rate µ.The PWM entries contribute additively to the mutation motif.
- Scenario enumeration: IGoR aligns each read to candidate germline genes, retains alignments above a configurable threshold, and constructs scenarios by selecting genes and deletion lengths.The remaining sequence between trimmed germline segments defines insertion possibilities.
- Scenario enumeration: Scenario likelihood is computed as Pscenario = Precomb × Perr, and candidate scenarios are ranked in decreasing probability.Perr represents sequencing-error or hypermutation probability under the applicable model.
- Generation probability: IGoR estimates Pgen by summing recombination probabilities over scenarios and averaging across possible unmutated sequences weighted by posterior probabilities.An alternative approximation uses the most likely pre-mutation sequence.
- Scenario enumeration: Computation is shortened by traversing a hierarchical decision tree and discarding branches whose upper-bounded probability contribution falls below a threshold.Only plausible scenarios are ultimately listed.
Learning algorithm
The learning algorithm estimates recombination and error or hypermutation parameters from sequence data, while evaluation uses divergences and targeted statistics to compare inferred, synthetic, and real repertoires.
- Learning algorithm: Expectation-Maximization infers model parameters from unique sequences by iteratively weighting possible scenarios and maximizing their weighted log-likelihood.The procedure repeats parameter updates until convergence.
- Evaluation: Kullback-Leibler divergence compares inferred parameters with known parameters from synthetic data across recombination scenarios.The divergence can be decomposed into additive contributions from scenario features.
- Evaluation: The radial distribution function measures correlations between hypermutations at nearby positions along BCR sequences.It is used to evaluate the occurrence of hypermutations at different separations.
- Evaluation: Double-D insertion frequency is estimated by aligning two non-overlapping D segments over at least 10 nucleotides and comparing synthetic with real sequencing data.The alignment uses the Smith-Waterman algorithm between the best V and J alignments.
A Model definitions
The model represents receptor generation as probabilistic recombination followed by sequence observation, while accommodating alternative chain types, dependencies, and read-level errors or hypermutations.
- A Model definitions: Each chain-specific model assigns probabilities to gene choices, insertions, and deletions while retaining only dependencies needed to explain observed correlations.Insertions are modeled as Markov chains with nonparametric length distributions.
- A Model definitions: For TCRα and BCR light chains, a scenario contains V and J choices, deletions, and VJ insertions; TCRβ and BCR heavy chains additionally include D-related variables.TRB gene usage is factorized as P(V, D, J) = P(V)P(D, J).
- A Model definitions: Bayesian networks encode configurable dependencies among recombination subprocesses, including gene choices, deletions, and insertions.The network is represented as a directed acyclic graph whose vertices label subprocesses.
- A Model definitions: The model parameterizes scenario and sequence probabilities, while read likelihoods incorporate sequencing errors or hypermutations after recombination.Read length affects generation probability, and insertion/deletion sequencing errors are ignored in the stated simplification.
B Expectation-maximization
Expectation-Maximization handles hidden recombination scenarios by alternating posterior weighting with parameter maximization, increasing likelihood until convergence.
- B Expectation-maximization: Expectation-Maximization is appropriate because multiple recombination and hypermutation scenarios can produce the same observed read.The scenario from which a read originated is generally unknown.
- B Expectation-maximization: The Expectation step computes posterior scenario probabilities, which weight the joint log-likelihood in the pseudo-log-likelihood objective.The objective separates recombination and error or hypermutation contributions.
- B Expectation-maximization: The Maximization step updates model components under normalization constraints by maximizing the pseudo-log-likelihood with fixed posterior weights.Lagrange multipliers impose normalization on the component probabilities.
- B Expectation-maximization: Iterative maximization of the pseudo-log-likelihood increases total likelihood, and the algorithm converges to a likelihood maximum.The stated inequality guarantees that the update improves total likelihood by at least the corresponding objective increase.
3 Optimizing the independent single nucleotide error model
IGoR models independent nucleotide errors and context-dependent hypermutations through separate likelihood objectives optimized within the EM framework.
- 3 Optimizing the independent single nucleotide error model: The independent single-nucleotide error model assigns mismatch probability r equally among the three alternative nucleotides at each error-prone position.The model treats nucleotide errors independently.
- 3 Optimizing the independent single nucleotide error model: Error-model optimization uses mismatch counts between reads and recombination products over potentially erroneous genomic nucleotides.The relevant length L depends on the scenario and read rather than equaling read length.
- 3 Optimizing the independent single nucleotide error model: The hypermutation model uses sequence context around each mutation, with position-weight parameters and an overall mutation rate included in θ.The context is a (2m+1)-mer centered on the mutation position.
- 3 Optimizing the independent single nucleotide error model: Hypermutation-model parameters are optimized at each EM maximization step using Newton’s method with backtracking line search.A constraint removes one parameter per position to resolve model degeneracy.
C Model entropy and DKL
The paper defines entropy as uncertainty in a stochastic recombination process and Kullback-Leibler divergence as dissimilarity between parameterized models. Cross-entropy provides the common computational basis for both quantities and can be decomposed by model component.
- Entropy measures uncertainty in outcomes generated by a stochastic process parameterized by θ.
- Kullback-Leibler divergence measures dissimilarity between two recombination distributions parameterized by θ1 and θ2.
- Entropy and Kullback-Leibler divergence can both be computed from cross-entropy between two parameterized distributions.
- For the considered models, cross-entropy can be divided into subparts corresponding to individual model components.
- Only processes directly or indirectly affecting component i need to be summed, rather than all possible recombination scenarios.These processes are its ancestors in the acyclic dependency graph.
2 Inserted nucleotides
The paper computes cross-entropy and generation-probability quantities by exploiting dependencies among recombination processes and by approximating ancestral sequences for noisy reads. These calculations remain tractable when ancestor sets are small, but mutated-sequence generation probabilities are intrinsically approximate.
- Insertion cross-entropy is computed separately for a fixed insertion length and then averaged over possible lengths.The inserted sequence is modeled through a stationary Markov-chain distribution.
- The insertion calculation accounts for processes affecting the insertion length or inserted sequence, excluding the insertion-length process itself.
- For a noisy or hypermutated read, IGoR estimates the generation probability of its unmutated ancestor using a posterior-weighted geometric average or the most likely recombination product.The posterior combines recombination and read-error probabilities.
- The distribution of log generation probabilities can be estimated from data and becomes exact as N →∞.
- IGoR uses sparse expectation maximization and branch pruning to reduce the number of scenarios explored.The parameter ε controls the sparsity approximation, with ε = 0 giving exhaustive exploration.
4 Generating synthetic sequences
The supplementary analyses evaluate synthetic-data recovery, convergence, computational cost, and comparisons across receptor types and software. They show accurate recovery under suitable sample sizes and error rates, while highlighting uncertainty and differences in assignment and inferred distributions.
- 4 Generating synthetic sequences: Synthetic sequences are generated by randomly sampling recombination scenarios and cutting the resulting sequences to mimic sequencing.
- 5 Comparison to other software: 3× faster on average, IGoR’s most-likely-scenario mode is faster than evaluating all scenarios per sequence.Restricting scenarios to deterministically assigned V and J genes is 6× faster, according to the caption.
- 5 Comparison to other software: IGoR converges toward the true insertion and deletion distributions on simulated data across sample sizes and error rates.
- 5 Comparison to other software: 26.5% correct predictions, IGoR outperforms MiXCR’s 9.8% and Partis’ 7.4% on sequences without palindromic insertions.
- 5 Comparison to other software: IGoR captures the physiological exclusion between TCR D2 and J1 in real data, whereas MiXCR does not.
- 5 Comparison to other software: BCR D-J usage does not show the clear coupling observed for TCRs.