Source-linked AI summary
Evolution of complex modular biological networks
Arend Hintze, Christoph Adami
TL;DR
The paper asks how modularity contributes to robustness and evolvability in biological networks and how it can be identified without a preconceived definition. It evolves artificial metabolic networks across environments with different predictability and analyzes their topology, information content, and genetic interactions. The networks acquire biological network properties, while synthetic lethal pairs tend to remain within modules and suppressor pairs tend to straddle them.
Problem
The study addresses how modularity supports robustness and evolvability and how modules can be identified using biological network evidence.
Method
The authors evolve artificial metabolic networks in environments with differing predictability and analyze topological, information-theoretic, and genetic-interaction measures.
Results
Evolved networks display scale-free and small-world properties, while synthetic lethal pairs usually lie within modules and compensatory pairs preferentially straddle modules.
Takeaways & Limitations
Combining network modularity tools with genetic interaction data provides an approach for studying modularity in the evolution and function of biological networks.
Takeaways & Limitations
The dynamic-environment result differs from earlier work because the environments and network types use different forms of change and representation.
Abstract
from arXiv · showhide
Biological networks have evolved to be highly functional within uncertain environments while remaining extremely adaptable. One of the main contributors to the robustness and evolvability of biological networks is believed to be their modularity of function, with modules defined as sets of genes that are strongly interconnected but whose function is separable from those of other modules. Here, we investigate the in silico evolution of modularity and robustness in complex artificial metabolic networks that encode an increasing amount of information about their environment while acquiring ubiquitous features of biological, social, and engineering networks, such as scale-free edge distribution, small-world property, and fault-tolerance. These networks evolve in environments that differ in their predictability, and allow us to study modularity from topological, information-theoretic, and gene-epistatic points of view using new tools that do not depend on any preconceived notion of modularity. We find that for our evolved complex networks as well as for the yeast protein-protein interaction network, synthetic lethal pairs consist mostly of redundant genes that lie close to each other and therefore within modules, while knockdown suppressor pairs are farther apart and often straddle modules, suggesting that knockdown rescue is mediated by alternative pathways or modules. The combination of network modularity tools together with genetic interaction data constitutes a powerful approach to study and dissect the role of modularity in the evolution and function of biological networks.
Introduction
The paper frames modularity as a higher-level organization that can make biological networks robust and evolvable, while noting that its definition and identification remain unsettled. It proposes evolving artificial metabolic networks and combining topological, information-theoretic, and genetic-interaction analyses to study modularity.
- Modules are functional building blocks whose functions are separable, enabling a synthetic, higher-level analysis of biological systems.
- The study evolves artificial metabolic networks and examines topological and information-theoretic modularity alongside simulated genetic interaction experiments.
- Modularity may allow evolutionary change with minimal disruption while supporting simultaneous optimization of robustness and evolvability.
- Biological modules can be identified using functional, evolutionary, or topological criteria, including shared regulation, clustered network connectivity, or co-inheritance.
- Genetic interactions such as synthetic lethality and dosage rescue provide information about cellular robustness and modularity.
- In silico evolution can generate complex networks with biological features and support knockdown or overexpression experiments that integrate functional, evolutionary, and topological information.
Results
The model evolves artificial cells whose genomes encode proteins acting on a chemically defined metabolite system in spatially structured environments. The resulting pathways can be represented through multiple network constructions, while phylogenetic depth tracks evolutionary progress.
- Artificial Chemistry: 608 valid molecules form the artificial chemistry, with reactions generated by cleavage that preserves atomic content.
- Artificial Chemistry: 5,020,279 theoretically possible cleavage reactions actually lead to valid molecules.
- Organisms: Cells inhabit a 2D chemostat where 53 precursor molecules diffuse from replenished sources, while cellular products are removed each update.
- Organisms: Cell division requires producing sufficient metabolites by importing precursors and catalyzing reactions with specific transporter and enzymatic proteins.
- Organisms: Genomes use a four-symbol alphabet and evolve with fitness-proportional selection, point mutation, and gene duplication.
- Environments: Three environments vary in precursor availability, including static and changing source locations, to model predictable and unpredictable conditions.
- Network representations: The same pathway can be represented as functional, metabolic, or protein-protein interaction graphs, each producing different topological properties.
- Phylogenetic Depth: Phylogenetic depth measures descent from the ancestral genome and serves as a proxy for generations elapsed in a run.
Network Evolution
Evolution produced increasingly complex metabolic networks with small-world, scale-free, hub-based, and modular organization, while robustness and modularity depended on environmental predictability. Genetic interactions further revealed that synthetic lethality is usually within modules, whereas compensatory interactions often span modules.
- Pathways became increasingly complex through gene duplication, divergence, pathway combination, and new precursor-import pathways, forming loops and multiple interconnections.
- Functional and metabolic representations of evolved networks were approximately scale-free, whereas their protein-protein representation had an exponential degree distribution.
- Average geodesic distances remained short as networks grew, producing small-world connectivity consistent with metabolic networks.
- Removing hubs caused a sharp path-length breakdown, whereas removing random nodes increased path length smoothly until near network collapse.The hub-removal breakdown occurred at about 200 hubs removed in the depicted functional network.
- Modularity increased over evolutionary time, but networks evolving in dynamic environments were generally less modular because precursor-production pathways connected major metabolic pathways.Dynamic environments also required more complex pathways for reliable function, slowing network evolution.
- Node-removal robustness barely declined with increasing fitness, while environmental robustness declined in static and quasi-static environments but stayed nearly constant in dynamic environments.
- Synthetic lethal pairs usually remained within modules, whereas compensatory pairs preferentially straddled modules across modularity definitions.The evolved metabolic networks differed from yeast protein-interaction networks in connectivity and in the nature of the compensatory and synthetic-lethal interactions.
Methods
The study evolves artificial-cell genomes and metabolic networks whose encoded proteins import, export, and catalyze reactions, then evaluates fitness, information, modularity, and robustness.
- Genome representation: Artificial genomes encode protein type, expression level, reaction or molecule specificity, and affinity using four nucleotide symbols.Genes begin with four zeros and may overlap; mutations in specificity regions remain mapped to legal reactions.
- Protein interactions: Protein affinity is computed from four active-site domains, each compared with one molecule involved in a reaction.The affinity score D(M,P) equals 1−S(M,P), where S is a similarity score, and perfect complementarity gives the highest affinity.
- Network dynamics: Networks are evaluated with discretized metabolic-rate equations that update molecule concentrations from reaction connectivity, flux, and affinities.The model includes diffusion of precursor molecules in a two-dimensional chemostat, with concentration determined by distance from defined sources.
- Fitness: Cells metabolize 608 possible molecules, with the first 53 designated precursors and the remaining 555 treated as increasingly complex metabolites.Fitness aggregates synthesized metabolites, and only metabolites reaching non-vanishing abundance contribute to the product.
- Fitness: Fitness is context dependent because precursor concentrations near cells affect performance, while its multiplicative form makes pathway discovery beneficial by a constant percentage.The study plots logarithmic fitness because it converts the multiplicative measure into an additive quantity.
- Evolutionary process: Evolution uses Wright–Fisher fitness-proportional selection with Poisson point mutations, capped mutation counts, adjacent duplication, and deletion, without recombination.Organisms are protected from death and cannot replicate until they are at least eight updates old.
- Network analysis: Modularity is quantified with information-bottleneck clustering, while information content is estimated from sequence length and entropy.The clustering assigns network nodes to descriptions by balancing graph-description simplicity against retained relevance; components smaller than five nodes are excluded from average modularity.
Supporting Information
Supporting analyses characterize evolved networks through degree and path-length distributions, robustness under removals, and spatial separation of genetic-interaction pairs.
- Network topology: Molecule participation follows an approximately power-law distribution with exponent λ≈2.23 across 80 dynamic-environment runs.The reported fit has r^2=0.88, with standard-error bars and threshold-based variable bin sizes.
- Network topology: Mean path length is compared for metabolic and protein-protein annotations across static, quasi-static, and dynamic environments.The supporting figure tracks this quantity along the evolutionary sequence shown in the main study.
- Robustness: Robustness analyses measure normalized log-fitness decline after precursor or node removal along the line of descent.The figure distinguishes ancestral depth from later evolutionary positions and separates precursor-removal from node-removal perturbations.
- Genetic interactions: Gene-pair separation is analyzed by interaction class, contrasting synthetic lethal, dosage-rescue or knockdown-suppressor, and random pairs.Additional analyses examine distance distributions and whether high-betweenness-node removal separates suppressor pairs.
- Environmental effects: Precursor-producing genes are compared between dynamic and static environments as the fraction involved in producing the 53 precursor molecules.The comparison is presented as a red-versus-green environmental analysis.