Source-linked AI summary

SwarmWorld: Stigmergic technological evolution in societies of language-model agents

Subhadeep Pal, Fiona Y. Wang, Markus J. Buehler

arXiv:2608.26081v1cs.AIcond-mat.mtrl-scics.CL

TL;DR

It remains unclear whether decentralized language-model agents can build functional technologies and outperform independent search. SwarmWorld studies initially homogeneous agents in a persistent shared world and finds broader, more resilient technological portfolios than isolated search, while isolated search remains competitive for the strongest artifact.

  • Problem

    It remains unclear how interaction changes technology discovery, accumulation, inheritance, and robustness relative to parallel independent search.

  • Method

    SwarmWorld places initially homogeneous agents in a persistent, materially constrained world where they explore, construct artifacts, and author executable controllers evaluated by a deterministic simulator.

  • Results

    Shared societies produced broader, more resilient technological portfolios than matched best-of-N isolated search and self-organized distinct exploratory and artifact-centered behaviors.

  • Takeaways & Limitations

    Physical stigmergy alone supported capable decentralized coordination, while cultural interaction fostered persistent technological ecologies without universally superior individual inventions.

  • Takeaways & Limitations

    Cultural benefits depended on the measured outcome and timescale, with no persistent full-culture advantage on held-out resilience.

Abstract

from arXiv · show

Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions to accumulate into durable social organization. Language-model agents offer a new substrate for this process, yet most multi-agent systems rely on direct conversation, predefined roles, or centralized workflows. It remains unclear whether decentralized agents can build functional technologies and outperform independent search. Here, initially homogeneous LLM agents in SwarmWorld self-organize without assigned roles or recipes into evolving technological societies. Agents explore a spatial environment, process resources, test materials, construct persistent artifacts, and write executable controllers evaluated by a deterministic simulator under unseen disturbances after the agents are removed. SwarmWorld splits cognition from consequence: agents propose architectures and controllers within fixed action and material schemas, while the simulated world determines function. Shared societies develop broader, more resilient technological portfolios than a strong best-of-N isolated-search baseline, although isolated search remains competitive for the strongest artifact. Agents differentiate into exploration, construction, maintenance, and coordination behaviors, transitioning as the world matures. Technologies accumulate through collaborative construction, executable inheritance, and persistent agent-artifact networks, with most reuse beginning through physical observation rather than communication. Explicit cultural mechanisms amplify collaboration and organization, but functional benefits depend on outcome and timescale. Physical stigmergy alone supports capable societies, while interaction drives persistent technological ecologies rather than universally superior individual inventions.

1 Introduction

SwarmWorld tests whether initially homogeneous, decentralized LLM agents can construct a cumulative shared substrate rather than merely search in parallel. It evaluates persistent executable technologies independently of agent claims against matched isolated search, finding broader, more resilient portfolios and self-organized differentiation in shared worlds.

  • Motivation: Decentralized collective behavior can amplify local feedback, adapt task allocation, and preserve environmental changes that alter later agents’ information and opportunities.These mechanisms motivate testing whether artificial collectives can build a shared, cumulative substrate for future action.
  • Research gap: SwarmWorld addresses a gap by combining initially equivalent agents, a persistent shared world, independently evaluated executable technologies, and a matched isolated-search baseline.The design tests whether interaction changes capability rather than merely increasing the number of samples.
  • SwarmWorld: Initially homogeneous LLM agents explore, process resources, test materials, construct persistent artifacts, and author executable controllers without assigned roles, recipes, or a technology catalog.The environment separates agent proposals from consequences determined by the world’s fixed action, material, and simulation constraints.
  • Evaluation: Controlled ablations isolate communication, cross-agent program inheritance, and physical stigmergy, while societies are compared with endpoint-wise best-of-N matched isolated agents.The comparison gives swarm advantage a falsifiable criterion: exceeding the capability achieved by the same computational population through independent search.
  • Findings: Across populations of 50–200 agents and complementary long-horizon experiments, shared worlds produced broader, more resilient technological portfolios and self-organized differentiation.Agents differentiated into behaviors associated with exploration, construction, maintenance, and coordination as worlds matured.

2 Results and Discussion

SwarmWorld shows that shared physical worlds produce broader, more resilient technological portfolios than isolated search, but cultural mechanisms yield endpoint- and timescale-dependent benefits rather than universal swarm superiority. Agents differentiate behaviorally and accumulate technologies through construction, executable inheritance, spatial coordination, and reusable knowledge lineages.

  • Experimental design: Two paired studies tested population scaling across four mechanism-resolved conditions and long-horizon development against an endpoint-wise best-of-100 independent-search envelope.The scaling study ran 800 ticks at N = 50, 100, and 200; the long-horizon study followed N = 100 societies for 3,200 ticks.
  • Portfolio outcomes: Shared worlds consistently produced broader, more resilient portfolios than isolated search, although isolated search could retain the strongest single artifact.No-explicit-culture societies sometimes outperformed full culture, so the main result was not a universal swarm advantage.
  • Population scaling: Population scaling was mechanism-dependent: at N = 200, no explicit culture produced the largest paired discovery gain, +0.069.Discovery-frontier AUC generally increased with population, but condition rankings changed across N; all three shared-world conditions exceeded the independent envelope at N = 100.
  • Population scaling: At N = 200, no explicit culture reached a mean paired gain of six validated inventions, while positive portfolio-resilience effects appeared throughout shared-world cells.An invention required tested materials, a complete design, an installed agent-authored program, threshold performance, and behavioral novelty.
  • Behavioral differentiation: Agents differentiated without role prompts: artifact-centered fractions at N = 200 averaged 27% under full culture, 20% without explicit culture, and 17% without communication.Artifact-centered work combined proximity, construction, control, and cultural coordination, whereas the other phenotype emphasized broader movement and lower artifact engagement.
  • Cultural and spatial organization: Technological accumulation involved multi-agent construction, executable program forking, spatial hubs, and reusable knowledge lineages linking observations and programs to downstream artifacts.Under full culture, 67%, 76%, and 56% of artifacts at N = 50, 100, and 200 recorded contributions from more than one agent; artifact-contact AUC at N = 200 was 0.31 under full culture, 0.14 without explicit culture, and 0.11 without communication.
  • Long-horizon dynamics: Over 3,200 ticks, full culture overtook no explicit culture in mean best-artifact performance by tick 800 and in portfolio resilience and cumulative artifact count near tick 1,600, but never in validated invention count.Resilience changed sign across checkpoints and was effectively tied at tick 3,200, showing that cultural benefits depend on the endpoint and timescale.

3 Conclusion

SwarmWorld shows a bounded swarm advantage: interaction improves resilient technological ecologies more reliably than isolated search for a single record-setting artifact. Societies self-organize through physical stigmergy, persistent artifacts, executable lineages, and modular relationships, while cultural benefits depend on outcome and timescale.

  • 3 Conclusion: Shared worlds consistently improved held-out resilience, portfolio resilience, and validated inventions relative to the endpoint-wise independent-search envelope.These gains appeared across the 800-tick scaling study, whereas isolated search remained competitive for the strongest artifact.
  • 3 Conclusion: Agents differentiated into artifact-centered and mobile-exploration phenotypes without assigned roles, while explicit culture increased the artifact-centered long-horizon fraction by 21.8 percentage points.Explicit culture also reduced movement, increased regional crowding, and sustained growth in executable lineage depth.
  • 3 Conclusion: The shared physical world served as external memory and transmission medium, turning movement into observation, construction, installation, and coordination around persistent sites.Recorded provenance linked evidence and programs to downstream artifacts, while accumulated modifications reshaped the search space for subsequent activity.
  • 3 Conclusion: Mature societies formed persistent, modular technological neighborhoods with local hubs and cross-community connectors rather than global synchronization or independent agents.Agent-artifact ties densified superlinearly, full culture produced roughly twice as many cumulative ties, and relationship reuse rose in both shared worlds.
  • 3 Conclusion: Cultural mechanisms produced outcome-dependent benefits: best-artifact performance favored full culture by about tick 800, portfolio resilience and artifact production crossed near tick 1,600, invention count never crossed, and held-out resilience showed no persistent advantage.Artifact stigmergy preserved high capability even without explicit culture.
  • 3 Conclusion: The evidence is bounded by four matched world seeds per condition, one model and prompting configuration, and simulator-defined material and environmental functions.Technology portraits visualize mechanisms, while the knockout assay measures graph topology rather than physical recovery after agents disappear from a running world.

4 Materials and Methods

SwarmWorld separates LLM cognition from deterministic physical execution: agents produce validated local plans, while the simulator enforces constraints, advances the world, and runs persistent artifact programs. Frozen portfolios are then evaluated agent-free under paired unseen disturbances to test whether accumulated technologies remain functional.

  • Agent control: Agents receive local observations, private memory, and condition-permitted shared records, then return schema-validated state updates and action plans executed through queued simulator actions.Runtime validation enforces location, ownership, empirical grounding, capacity, and condition permissions; invalid output is logged.
  • Simulation: The deterministic simulator checks spatial, material, energetic, and treatment-specific constraints, advances environmental resources, executes installed artifact programs, and records outcomes as experience.Communication, publication, inheritance, and artifact visibility vary by experimental condition, while roles and recipes are not assigned.
  • Agent-free evaluation: At checkpoints, frozen worlds and technological portfolios are copied, agents are removed, and deterministic physics plus installed programs operate under paired previously unseen disturbance schedules.This agent-free evaluation measures whether accumulated technology remains functional without continued LLM intervention.
  • World generation: Worlds are deterministic for fixed seed and parameters, with validated layouts satisfying structural invariants such as walkable area, required resources, nonoverlapping facilities, and reachable interaction regions.Scenario packages can replace or extend world definitions through validated YAML mapped to stable typed interfaces without loading package-authored executable code.
  • Study design: The 800-tick study covered 4 interaction conditions, 3 population sizes (N = 50, 100, 200), and 4 matched seeds, yielding 48 episodes; the 3,200-tick study used N = 100 and 4 new seeds.Frozen portfolios were evaluated under 8 paired held-out disturbance schedules, with comparisons within matched population–seed blocks using deterministic paired bootstrap intervals and exact sign-flip tests for the long-horizon study.

Data and code availability

The authors make the paper’s code, protocols, outputs, analysis and manuscript sources publicly available, alongside additional experimental data.

  • Data and code availability: Code, exact prompt manifests, versioned protocols, raw model outputs, analysis scripts, figure-generation code, and manuscript sources are available on GitHub, with additional experimental data on Hugging Face.Repositories: https://github.com/lamm-mit/SwarmWorld and https://huggingface.co/datasets/lamm-mit/swarmworld-data.

Supplementary Information

The supplementary information identifies the paper, its authors, and their Massachusetts Institute of Technology affiliations.

  • The paper is titled “SwarmWorld: Stigmergic technological evolution in societies of language-model agents.”
  • The authors are Subhadeep Pal, Fiona Y. Wang, and Markus J. Buehler.
  • The authors are affiliated with departments and centers at the Massachusetts Institute of Technology in Cambridge, Massachusetts, USA.

S1 Glossary of key terms

The glossary defines implementation-aligned measures for isolated-search performance, artifact discovery and validation, agent behavior, and technological-network organization. It distinguishes outcome-based function from structural properties such as lineage, modularity, nestedness, scaling, and reuse.

  • Performance and validation: The independent-search envelope takes the best isolated-agent result separately for each endpoint and checkpoint.The maximizing agent may differ across endpoints and checkpoints, so one solo agent need not win every comparison.
  • Performance and validation: Discovery-frontier AUC measures the time-averaged running maximum of artifact performance, rewarding strong inventions that appear early.The frontier is F(t) = maxτ≤t p(τ), with normalized area computed by trapezoidal integration over recorded samples.
  • Performance and validation: Balanced coverage scores the best active artifact for each service dimension while penalizing neglected needs.The score combines mean service coverage with the minimum-to-mean balance ratio; artifact count alone does not increase discovery-state scores.
  • Performance and validation: Validated invention requires tested materials, complete design fields, an executable agent-authored program, sufficient performance, and behavioral novelty.The frozen technology is tested under new disturbances after agents are removed, measuring continued coverage of multiple needs.
  • Social and technological organization: Agent-artifact organization is characterized through observed behavioral phenotypes, recorded lineage and reach, community structure, nestedness, scaling, and relationship reuse.These measures capture trajectory-derived behavior, inherited programs and contributions, within- versus cross-community connectivity, Ecum ∝V α scaling, and persistence of prior ties.

S2 Supplementary Methods · S2.1 Agent model, communication, and model-to-world interface · S2.2 World representation and constructing custom worlds

The supplementary methods define how adoption and network persistence are measured, how homogeneous agents observe, plan, and act safely, and how deterministic worlds are generated and controlled across experiments.

  • S2 Supplementary Methods: Adoption requires someone other than the inventor to engage with a technology, and delay measures how quickly that reuse occurs.Borrowing and using a tool counts as adoption, whereas merely helping build it does not.
  • S2 Supplementary Methods: The network-persistence score measures how much of the recorded network remains connected as agents disappear, not physical recovery or adaptation.An internet topology illustrates that tolerance of random failures does not imply resilience to targeted hub removal.
  • S2.1 Agent model, communication, and model-to-world interface: Agents were homogeneous across conditions, using gpt-5.6-luna at temperature 0.7 with shared prompts, schemas, capabilities, inventory, and memory budgets.The configuration allowed 4,096 output tokens, 12 planned actions, a 60,000-character retrieved-context budget, and 64 private memory records.
  • S2.1 Agent model, communication, and model-to-world interface: Agents receive local semantic observations, sparse empirical maps of seen locations, and private memory containing outcomes, notebook evidence, recipes, and research state.Observations include terrain, resources, facilities, agents, artifacts, environmental measurements, inventory, pending microbatches, and action affordances.
  • S2.1 Agent model, communication, and model-to-world interface: Each response must be one validated JSON object containing a research-state update and plan, while runtime checks enforce state-dependent preconditions.Invalid outputs become logged safe waits and never reach evaluation, execution, shells, or the artifact virtual machine.
  • S2.2 World representation and constructing custom worlds: Each simulation world is a rectangular two-dimensional lattice with typed terrain, resource, facility, and continuous environmental-field layers.World generation is deterministic for fixed random seed s and parameter specification θ, with candidate layouts checked against structural invariants.
  • S2.2 World representation and constructing custom worlds: World parameters θ specify dimensions, spatial-generation rules, resource capacities and renewal, environmental dynamics, facility constraints, and disturbance processes.Validation requires sufficient walkable area, required resource classes, nonoverlapping facilities, and reachable interaction regions.
  • S2.2 World representation and constructing custom worlds: Experiments fixed the lattice at 72 × 54 cells as population increased, so population altered agent density rather than available area.Shared discovery seeds produced identical worlds and disturbances, while deterministic nested-permutation starting positions preserved the first N positions across population sizes.

S2.2.1 State representation and reproducibility · S2.2.2 Declarative geometry, resources, and facilities · S2.2.3 Environmental fields and disturbances

SwarmWorld represents reproducible worlds through layered state, declarative scenario packages, and persistent traces of resolved configurations. Resources, facilities, geometry, environmental fields, and held-out disturbances are specified through structured rules that preserve replay while varying material settings and evaluation conditions.

  • S2.2.1 State representation and reproducibility: The authoritative state combines integer terrain, resource, and facility layers with floating-point mass and capacity layers and named continuous environmental fields.Run-level parameters include grid dimensions, disturbance interval and intensity, resource and field capacities, and an optional scenario package path.
  • S2.2.1 State representation and reproducibility: Every trace records the resolved configuration, engine revision, realized generator manifest, scenario identifier and version, and a SHA-256 hash over the scenario.
  • S2.2.2 Declarative geometry, resources, and facilities: Scenario packages map stable internal slots to public identifiers for nine terrains, eight nonempty resources, six nonempty facilities, and ten process operations.This preserves observation and replay while allowing one simulation engine to express different materials settings.
  • S2.2.2 Declarative geometry, resources, and facilities: Grid coordinates are normalized as ξ = x/ max(1, w −1) and η = y/ max(1, h −1), while geometry assigns base terrain before applying ordered feature masks.Later geometry features can override earlier assignments according to the declarative ordering.
  • S2.2.2 Declarative geometry, resources, and facilities: Resource ledgers initialize cell mass from capacity and configured fractions, while harvesting removes mass and overlapping deposits replace earlier values.The supplied specification also defines capacity and initial-mass expressions and optional renewal.
  • S2.2.3 Environmental fields and disturbances: Each environmental field declares a numerical range, diffusion and decay coefficients, per-terrain sources, and an initial condition assembled from constants, gradients, Gaussian terms, terrain offsets, and optional seeded noise.
  • S2.2.3 Environmental fields and disturbances: Cyclic fields are overwritten after updates by spatially uniform sinusoidal values, while version 1 requires temperature, water_availability, ground_stability, toxic_gas, and solar.Additional fields remain available in observations, snapshots, analysis, and rendering; updates use edge-value padding for the four-neighbor Laplacian.
  • S2.2.3 Environmental fields and disturbances: Held-out evaluation resamples disturbance centers and order from an evaluation seed while leaving the frozen technological state unchanged.Disturbances apply field-specific deltas and may trigger thresholded terrain transformations.

S2.2.4 Package boundary and example

Scenario packages are YAML-defined and bounded from executable code and physics-changing presentation logic. The example package organizes terrain, geometry, fields, resources, operations, facilities, artifacts, missions, disturbances, rendering, and analysis through referenced files.

  • Package boundary: Scenario packages cannot import Python, execute shell commands, or modify the active package during an episode.They may define terrain, field-overlay, and alias properties, but the current implementation does not load package-supplied meshes, textures, or shader source.
  • Example package: The example_domain package manifest uses format_version 1 and references separate YAML files for the scenario components.Referenced documents include terrains, geometry, fields, resources, operations, facilities, artifacts, missions, disturbances, rendering, and analysis.
  • Example package: The example geometry defines normalized-coordinate layers with PLAIN base terrain and ellipse, ridge, and rectangle features.The listed features assign CORRIDOR and WORKSPACE terrain to two shapes, while the rectangle specifies bounds [0.42, 0.48, 0.60, 0.66].
  • Example package: The example resource deposit has capacity 3.0, capacity_variation 0.15, initial_fill [0.65, 1.0], and regrowth 0.001.The example fields define Temperature with range [0.0, 1.0], diffusion 0.04, decay 0.001, constant initial 0.25, noise 0.01, and a radial component.

S2.3 Open-ended materials invention

Agents invent bioinspired material systems through locally available matter, typed processing, fabrication, and testing in an open-ended MATERIAL_SYSTEM space. Function is determined by deterministic material properties and closed physical flux accounting rather than textual descriptions or predefined artifact catalogs.

  • S2.3 Open-ended materials invention: Agents harvest local matter, formulate typed recipes, operate distributed workstations, fabricate private microbatches, and test them for deterministic normalized properties.Testing reveals properties only after fabrication, and feedstock can move between personal inventories and shared depots when allowed.
  • S2.3 Open-ended materials invention: The evaluator maps composition, processing order, hydration, porosity, alignment, crosslinking, and quality to normalized material properties without exposing its equation or global reward.Agents receive only outcomes from their own admissible operations and tests.
  • S2.3 Open-ended materials invention: A single generic MATERIAL_SYSTEM artifact class replaces catalogs of membranes, lattices, scaffolds, or preferred biological analogies.Artifact specifications include agent-authored names, claimed functions, architectures, inspirations, predicted effects, geometry, tested batches, and optional controllers; text does not change function.
  • S2.3 Open-ended materials invention: Closed artifact fluxes enforce source–sink accounting for water, contamination, embodied reserves, recharge, regrowth, disturbances, metabolism, and artifact transfers.Water removal equals storage input, remediation cannot exceed existing contamination, and growth, healing, repair, and nutrient release consume bounded embodied reserves.

S2.4 Persistent executable artifacts · S2.5 Interaction channels and recorded provenance

SwarmWorld makes persistent artifacts executable, identifiable, and inheritable under deterministic constraints, while interaction channels and provenance records connect agents’ messages, materials, artifacts, and programs. Physical stigmergy remains possible without symbolic communication, and ablations remove capabilities from both the model schema and execution engine.

  • S2.4 Persistent executable artifacts: Agents install persistent straight-line controllers with 1–64 instructions over 16 floating-point registers and sensors exposing local environmental, artifact, storage, and material properties.Available operations include constants, copying, arithmetic, extrema, and comparisons; actuators are capability-scoped.
  • S2.4 Persistent executable artifacts: Controllers have no jumps, loops, calls, imports, dynamic allocation, network access, file access, or code-interpreted strings.Registers are clipped to [−4, 4], extensive actuators are capped at 0.05 normalized units per tick, and programs execute on later simulator ticks, including agent-free evaluation.
  • S2.4 Persistent executable artifacts: SHA-256 of canonical instruction content determines each program identifier, while the registry records exact parent–child identifiers, authors, installations, and instruction differences.When forking is enabled, agents may use only programs they authored, observed, were taught, or inherited, and children must change at least one instruction.
  • S2.5 Interaction channels and recorded provenance: Interaction channels support broadcast and addressed messages, append-only publication, teaching, trade, task claims, shared depots, design composition, program reuse, and program forking.Messages and evidence receive stable identifiers, with reply and fulfillment identifiers linking requests to later successful actions.
  • S2.5 Interaction channels and recorded provenance: Incoming messages wait until the recipient’s next fixed macroturn, so communication does not purchase extra model calls.This timing rule preserves the fixed model-call structure while allowing message-based interaction.
  • S2.5 Interaction channels and recorded provenance: Provenance separates epistemic from physical contribution by recording recipes, tests, contributors, feedstocks, causal evidence, artifact ancestry, specifications, program history, authorship, and parentage.Citations cannot fabricate matter, and co-location alone does not count as intellectual or physical contribution.
  • S2.5 Interaction channels and recorded provenance: Physical stigmergy operates without symbolic channels because harvesting, deposition, visible artifacts, resource gradients, stored matter, damage, and services alter what later agents encounter.Experimental ablations remove capabilities from both the advertised model schema and executable engine contract rather than merely instructing models not to use them.

S2.6 Experimental conditions · S2.7 Completed 800-tick population study

The study compared four decentralized experimental conditions with independent search across matched population episodes, then evaluated frozen societies under unseen disturbances after removing agents. Its design standardized decision opportunities and execution settings while preserving distinct communication and culture interventions, with documented provenance caveats.

  • S2.6 Experimental conditions: Four conditions distinguished no communication from no explicit culture, because physical executable-program observation and descent remained available only in the former.No explicit culture also removed access to explicit programs, skills, authored text, and mutation parents while retaining physical-phenotype and environmental-consequence evidence.
  • S2.6 Experimental conditions: Independent search was evaluated as an endpoint-wise envelope rather than one agent’s trajectory, with winners varying across discovery AUC, final performance, resilience, invention count, and held-out evaluation.The isolated winner could also change between temporal checkpoints.
  • S2.6 Experimental conditions: Every isolated member received the same fixed decision opportunities as its corresponding shared-world agent, with configured model-call or action budgets partitioned across members.This preserves a matched comparison between shared-world and isolated-search conditions.
  • S2.7 Completed 800-tick population study: The primary study crossed four conditions with N ∈{50, 100, 200} and discovery seeds 3201–3204, producing 48 matched condition–population–seed episodes over 800 ticks.At a macroturn interval of 50, each agent received 16 scheduled decisions.
  • S2.7 Completed 800-tick population study: Population size changed density within a fixed 72 × 54 world, while all cells shared the same model, prompt, action schema, configuration hash, engine revision 9, and held-out disturbance seeds 9201–9208.Shared and isolated members also had identical phases.
  • S2.7 Completed 800-tick population study: After discovery, each complete final state was frozen, copied eight times, and advanced for 288 physics ticks without agents, model requests, or agent actions.Deterministic field dynamics and installed artifact programs continued under unseen schedules varying drought, contamination, damage, and resource variation.
  • S2.7 Completed 800-tick population study: Averaging eight unseen schedules produced one held-out observation per discovery seed, while artifact knockouts and portfolio assays used isolated copies without feeding information back into discovery.These procedures separated discovery from post-discovery evaluation.
  • S2.7 Completed 800-tick population study: The matrix combined eight recorded invocations with identical pooled configuration, prompt, and action-schema hashes, but retained dirty-worktree provenance caveats and one N = 200 infrastructure deviation.Manifests identified engine revision 9 and commits 83b9e7a, ab68bab, 76fa204, a13caee, and 738cef6.

S2.8 Completed 3,200-tick long-horizon study … S2.11 Statistical analysis

The study used a separately analyzed, four-block 3,200-tick long-horizon design with frozen agent-free evaluations and endpoints distinguishing discovery history, final function, portfolio resilience, and validated invention mechanisms. Behavioral, network, lineage, and role analyses were reconstructed from recorded events, while inference treated independently generated world seeds as the experimental units and emphasized paired effects over dichotomous significance.

  • S2.8 Completed 3,200-tick long-horizon study: The long-horizon study comprised 12 episodes in four matched seed blocks, each running 3,200 discovery ticks with 6,400 scheduled model decisions.Frozen, agent-free copies were evaluated at ticks 400, 800, 1,600, 2,400, and 3,200 under held-out seeds 9201–9208.
  • S2.8 Completed 3,200-tick long-horizon study: The long-horizon study was analyzed separately from the 800-tick matrix and tested when explicit cultural channels change technological outcomes.Its longer horizon and new seed block were not additional replicates for population-scaling studies.
  • S2.9 Endpoints and frozen evaluation: Discovery-frontier AUC measures time-normalized area under the immutable running maximum, rewarding early discovery and sustained improvement, whereas final-state performance is reported separately.Replacing a controller can reduce current function without erasing a historical discovery.
  • S2.9 Endpoints and frozen evaluation: Portfolio resilience combines mean service coverage with balance through the factor 0.5 + 0.5b, where b is the minimum-to-mean coverage ratio.Held-out resilience AUC averages this balanced current-service measure over 288 evaluation ticks and does not directly reward artifact count.
  • S2.9 Endpoints and frozen evaluation: Validated inventions required thresholded recipes, named claims and architecture, agent-authored programs, above-threshold lifetime performance, and behavioral novelty.Supporting endpoints included construction, program forking, lineage, adoption, causal closure, and material or program provenance.
  • S2.10 Behavioral, lineage, and network analyses: Behavioral organization and role transitions were inferred from unlabeled trajectory, activity, movement, coverage, artifact, and agent-exposure features rather than assigned occupations.Temporal analysis used nonoverlapping 200-tick windows, with a separate physical/task-only sensitivity excluding cultural features.
  • S2.10 Behavioral, lineage, and network analyses: Agent–artifact networks were reconstructed from recorded social, construction, program, repair, and executable-descent events, with Louvain communities and null-filtered coordination projections.Simplified network backbones were used only for visualization, while projections retained overlaps exceeding a degree-conditioned hypergeometric null after Benjamini–Hochberg correction.
  • S2.11 Statistical analysis: The independently generated world seed was the unit of inference, with nested agents, ticks, artifacts, forks, windows, edges, and schedules never counted as independent replicates.Conditions were compared within matched seed and population, and held-out schedule sets were averaged before seed-level inference.

S2.12 Reproducibility and data release · S2.13 Performance-ranked selection of semantically distinct technologies

The paper releases a deterministic, hash-verified reproducibility package containing the engine, data, traces, analyses, and figure-generation tools. It selects semantically distinct technologies by ranking designs, excluding exact and near duplicates, enforcing cluster caps, and requiring feasible target counts.

  • S2.12 Reproducibility and data release: The source repository retains the engine, configurations, replay and analysis tools, frozen figure inputs, and deterministic journal-figure generators.Offline figure generation verifies source SHA-256 hashes and writes an output manifest; rebuilding figures does not call language-model or image-generation services.
  • S2.12 Reproducibility and data release: The data builder validates complete 48-episode 800-tick and 12-episode long-horizon matrices before staging a release.The release retains manifests, summaries, compressed traces, isolated-member traces, derived analyses, figures, and a file-level SHA-256 inventory.
  • S2.12 Reproducibility and data release: Standalone release scripts reproduce endpoint summaries, stream movement trajectories, summarize trace events and actions, and verify every released file.These scripts support independent checking of the released artifacts and analyses.
  • S2.13 Performance-ranked selection of semantically distinct technologies: Algorithm 2 ranks and selects technologies across all designs using a target count K, maximum similarity τ, cluster cap M, and embedding model E.Each technology is represented by semantic text covering its name, architecture, function, inspiration, materials, fabrication, output, design principles, and controller operations.
  • S2.13 Performance-ranked selection of semantically distinct technologies: Technologies are ordered by recorded performance after semantic text is embedded and normalized for comparison.The semantic representation integrates recorded technological and controller attributes before selection.
  • S2.13 Performance-ranked selection of semantically distinct technologies: The selector excludes exact semantic-text duplicates, rejects candidates from full clusters, and filters semantic near-duplicates whose similarity exceeds τ.Accepted technologies are added until |S| = K, while cluster counts and duplicate hashes are updated after each selection.
  • S2.13 Performance-ranked selection of semantically distinct technologies: If fewer than K technologies can be selected under the declared constraints, Algorithm 2 raises an error because the constraints are infeasible.Otherwise, it returns the selected representative set S.

S3 Agent communication and simulator-validated consequences

The trace shows communication supporting grounded information sharing, executable inheritance, durable teaching, and simulator-enforced material exchange. Agents verify observations locally, compare inherited programs under matched conditions, and treat teaching and trade as auditable transactions rather than unrestricted belief copying.

  • Trace scope: The representative N = 100, seed-3301 trajectory ran for 3,200 ticks and contained 2,914 delivered messages, 40 teaching events, 52 trades, 457 installations, and 389 artifacts.These examples illustrate distinct mechanisms rather than estimating communication frequency or average effect.
  • Grounded communication: Agents convert local observations into navigational leads, but independently inspect and harvest resources rather than copying another agent’s belief.A013’s request and inspection are followed by A000’s independently executed harvest at the reported location.
  • Executable inheritance: Agents exchange measurements, modify persistent controllers through parent-child program edges, and compare descendants with parents under matched local conditions.The comparison reports descendant performance 0.0969 versus parent-program performance 0.0851, while explicitly limiting interpretation to current local conditions.
  • Durable teaching: A TEACH action transfers a controlled comparison into a durable, auditable record with explicit evidence identifiers for later retrieval.The simulator creates teaching_0000016669 after the action succeeds, citing causal parents test_00000080 and test_00000052.
  • Transactional exchange: A TRADE action invokes simulator-enforced transfer of physical matter, moving 0.20 CHITIN into the recipient’s inventory before subsequent construction.The trade creates fulfillment_0000031213, but the episode does not establish that the transferred mass alone caused the artifact.
Loading 2608.26081v1…