Source-linked AI summary

Functional anatomy of Pythia-Herwig differences with Kolmogorov-Arnold networks

Arghya Chattopadhyay

arXiv:2608.15952v1hep-phcs.LG

TL;DR

Global classifier scores quantify Pythia–Herwig disagreement but do not show which observable structures carry it across generator stages. This paper uses additive KANs to decompose and transport observable responses, finding that the discrepancy shifts from multiplicity to mass and shape, with only some early-stage information persisting downstream.

  • Problem

    Global classifier scores quantify generator disagreement but provide limited evidence about which observable structures carry it and persist across event-generation stages.

  • Method

    The study uses additive KANs to decompose classifier-derived log density ratios into explicit one-dimensional observable responses that can be inspected, recomposed, and transported across generator stages.

  • Results

    The difference is dominated by multiplicity at shower level, shifts toward jet mass and shape after hadronization, and becomes mixed shape–multiplicity driven in the full-generator configuration.

  • Takeaways & Limitations

    The Pythia–Herwig difference changes with simulation stage, while shower-level multiplicity can persist downstream and shower-level shape need not.

  • Takeaways & Limitations

    The fitted classifier logit is only an approximation to the conditional log density ratio, so reweighting performance must be tested rather than assumed.

Abstract

from arXiv · show

Differences between high-energy event generators can arise at several stages of the collision simulation, from the hard scattering through parton showering and hadronization to the final event. These differences are usually summarized using observable distributions or global classifier scores. While these quantify the disagreement, they do not reveal which observable-level structures carry it or whether those structures persist through different stages of event generation. In this work, we formulate this problem as a staged functional analysis of generator-model differences. Following the same hard dijet events through Pythia and Herwig at shower-only, hadronized, and full-generator levels, we use an additive Kolmogorov-Arnold network (KAN) representation of the classifier-derived log density ratio to decompose the learned discrepancy into explicit one-dimensional observable responses that can be isolated, recomposed, and transported between generator stages. Within the same eight-observable jet representation, the Pythia-Herwig difference is driven mainly by multiplicity at shower level, shifts toward jet mass and shape after hadronization, and develops a mixed shape-multiplicity driven structure in the full-generator configuration. Transporting the individual shower-level functional components downstream shows that shower-level multiplicity information can retain its reweighting power, whereas the corresponding shape responses need not do so even though shape becomes important again at later stages. The jet-mass factors, meanwhile, are limited by poor statistical support. This KAN-based framework therefore provides a functional anatomy of generator-model dependence, exposing both persistent structures and support failures that are hidden inside a single global classifier-derived reweighting function.

1 Introduction

This work analyzes Pythia–Herwig differences across controlled shower-only, hadronized, and full-generator stages using an additive KAN decomposition over a common eight-observable dijet representation. The explicit observable responses enable stage-by-stage anatomy, recomposition, Shapley allocation, and transport of individual learned factors through the event-generation chain.

  • Motivation: Pythia and Herwig implement physically motivated but different perturbative and non-perturbative prescriptions, so their generator differences can probe modeling choices across event-generation stages.The stages include hard scattering, parton showering, hadronization, and underlying-event modeling such as multiparton interactions and beam remnants.
  • Controlled staged comparison: The study evolves the same hard pp →jj events through shower-only, shower-plus-hadronization, and full-generator configurations, while keeping the event selection and eight-observable representation fixed.The full-generator configuration includes MPI and the underlying event, enabling controlled localization of changes in observable structure.
  • KAN functional representation: An additive KAN represents classifier-derived density-ratio differences through explicit one-dimensional observable responses, making each factor inspectable, exportable, and exactly factorizable within the fitted model.This exposes observable-level structure that would otherwise remain inside a single black-box classifier output.
  • Factor transport: Individual learned factors are transported directly through the event-generation chain to test whether observable-level components retain their reweighting relevance between stages.The staged construction supports comparison of how observable content reorganizes as hadronization and later generator components are added.
  • Stage-dependent anatomy: The analysis compares additive and deeper KANs and uses functional decomposition, subset recomposition, and Shapley allocation to determine which observables carry the Pythia–Herwig difference at each stage.Analysis I studies the difference separately at shower-only, hadronized, and full-generator levels.

2 Hard process and generator stages

The analysis compares Pythia and Herwig from a shared hard-scattering event sample across shower-only, hadronized, and full-generator stages. Jets and observables are reconstructed independently at each stage, preserving a common event basis without implying uniquely matched jets or isolation of a single microscopic modeling ingredient.

  • Hard process and generator stages: Leading-order pp →jj events at √s = 13 TeV are passed as the same LHE records to Pythia 8.312 and Herwig 7.3.0.The common matrix-element event fixes the hard scattering, so the comparison begins with generator evolution rather than different hard events.
  • Hard process and generator stages: The three configurations separately enable showering alone, showering plus hadronization, or showering, hadronization, MPI, and underlying-event modeling.Stage A disables hadronization and MPI; Stage B adds hadronization while retaining MPI and underlying-event exclusions; Stage C uses the full configuration.
  • Interpretation and limitations: The staged comparison identifies where the Pythia-Herwig difference changes but does not isolate a single microscopic ingredient or compare hadronization models on identical showers.Stage transitions are controlled configurations sharing the hard-event identity, not uniquely matched stochastic shower histories or reconstructed jets.
  • Jet reconstruction: Jets are reconstructed afresh at each stage with anti-kT, R = 0.4, using shower partons at Stage A and stable visible particles with neutrinos removed at Stages B and C.FastJet is rerun independently for both generators, and leading and subleading jets are newly defined by reconstructed pT.
  • Observable representation: The eight-observable representation remains defined across all six reconstructed representations, avoiding stage-dependent event removal from the common hard-event cohort.The variables cover constituent counting, invariant mass, momentum sharing, and radial structure; multiplicity’s microscopic meaning changes when inputs switch from partons to hadrons.
  • Interpretation and limitations: The density ratio excludes differences in overall event yield, generator failures, and events outside common acceptance.It is therefore defined as a probe of physics differences within the shared accepted sample rather than of cross-section normalization or generator reliability.

3 Theoretical framework: KAN and functional reweighting

This section frames generator reweighting through KANs, whose trainable one-dimensional edge functions make the classifier-derived log density ratio accessible as observable-level components. For equal class priors, the Bayes-optimal classifier logit equals the Herwig-to-Pythia log density ratio, while finite-sample reweighting remains dependent on distributional support.

  • KAN architecture: KANs replace fixed MLP activations with trainable one-dimensional functions on network edges, implemented here with B-spline components and a smooth base function.The learned responses are exported as dense numerical tables that can be evaluated, removed, and recombined without retraining.
  • KAN architecture: The width- KAN gives an additive raw logit over the eight jet observables, while the deeper width- network permits combinations between observables.The additive architecture is chosen to expose the learned generator difference directly at the level of individual observables rather than to benchmark classification performance.
  • Classifier reweighting: For equal class priors, the Bayes-optimal classifier logit is exactly the logarithm of the Herwig-to-Pythia density ratio on the common cohort.The implementation trains directly on the raw logit, which approximates the log density ratio but need not equal the Bayes-optimal function for finite networks and samples.
  • Classifier reweighting: Reweighting requires sufficient overlap: source Pythia events cannot represent target Herwig regions absent from the source sample, and concentrated weights reduce statistical information.Support reliability is assessed separately through effective sample size and influence diagnostics rather than inferred from component allocations alone.
  • Interpretation: Observable-level Shapley allocations describe contributions within the fitted additive KAN, not causal fractions of shower, hadronization, or MPI physics.Physical interpretation instead comes from comparing observable structures and relative importance across the controlled generator stages.

4 Analysis I: stage-dependent anatomy of the generator difference

The Pythia–Herwig discrepancy reorganizes across generator stages: multiplicity dominates at shower level, mass and shape become important after hadronization, and mixed shape–multiplicity structure emerges in the full generator. Additive KAN responses preserve nearly all classifier information while exposing support limitations, especially for mass-based reweighting.

  • Functional decomposition: The additive KAN sacrifices only little of the observed generator information, enabling interpretation of individual one-dimensional responses within the eight-observable representation.This is an empirical property of the chosen representation, not evidence that the true density ratio factorizes or that inputs are uncorrelated.
  • Stage A: At stage A, constituent multiplicity dominates the generator difference, with lower multiplicities favoring Herwig and larger multiplicities favoring Pythia.The multiplicity response is supported by sizeable event populations, and the five largest component weights carry only a small fraction of total multiplicity weight.
  • Support diagnostics: Jet-mass responses grow in sparsely populated Pythia regions, causing a few events to dominate their weights and making standalone mass reweighting statistically fragile.The resulting effective sample size becomes very small.
  • Stage B: At stage B, the large shower-level multiplicity response disappears, while non-trivial dependence remains in jet-mass and shape observables.This stage reorganizes the difference after each generator applies its own shower and hadronization model; it is not a causal isolation of hadronization effects.
  • Stage C: At stage C, multiplicity dependence reappears alongside persistent mass and shape dependence, producing a shape-led structure with a substantial multiplicity component.Lower multiplicities contribute positively to the Herwig logit, whereas larger multiplicities contribute negatively.

5 Analysis II: transporting the shower-level functional response

Transporting frozen shower-level KAN responses tests whether their encoded information remains useful for reweighting later generator stages. Multiplicity responses retain transport power, whereas shape responses generally do not, and mass responses are unusable because of poor statistical support.

  • Transport construction: Transport evaluates each shower-level factor on the Pythia representation of a hard event and carries that unchanged to later stages.It tests persistence of encoded information associated with the same hard scattering, not persistence of a matched shower or jet trajectory.
  • Multiplicity transport: Leading-jet multiplicity provides the clearest supported transport direction, with Neff/N = 0.1080 and consistent direction across five stage A models.The held-out final-test sample confirms the same qualitative direction, although its separation reduction is weaker than on development data.
  • Multiplicity transport: Subleading-jet multiplicity also transports on development data, with Neff/N = 0.1408, and its downstream effect becomes relatively stronger at stage C.All five stage A models support this direction, suggesting multiplicity information is reorganized through later generator evolution.
  • Shape transport: Shape factors retain strong statistical support, with Neff/N ≃0.83-0.95, but none appreciably reduces downstream generator separation.The leading pT, leading-girth, subleading pT, and subleading-girth factors produce downstream AUCs that remain separated from 0.5, so the absence of transport is not caused by weight concentration.
  • Mass transport: Mass-factor transport is limited by statistical support because sparse source-stage tails create highly concentrated importance weights.In the representative fit, only 2 events accumulate 50% of the leading-mass weight and only 1 event accumulates 50% of the subleading-mass weight.

6 Conclusion and discussion

The paper introduces an additive-KAN framework that decomposes Pythia–Herwig differences into inspectable one-dimensional functional responses across generator stages. The analysis shows that these structures evolve through the simulation chain, with only some early-stage information remaining statistically usable downstream.

  • Methodological contribution: The KAN framework decomposes the learned Pythia–Herwig discrepancy into explicit one-dimensional responses that can be inspected, recomposed, stress-tested, and transported across generation stages.The comparison follows common hard dijet events through shower-only, hadronized, and full-generator configurations using eight observables.
  • Stage-dependent anatomy: The discrepancy changes substantially by stage: multiplicity dominates shower-only events, jet mass and shape dominate after hadronization, and full-generator differences become mixed shape–multiplicity structures.The overall separation drops sharply after hadronization, while multiplicity reappears as a substantial full-generator component.
  • Transport and persistence: Shower-level multiplicity factors retain statistically supported downstream reweighting power, including on the held-out final-test sample, whereas shape factors do not reduce downstream generator separation.The shape responses remain statistically well supported, but their stage-A factors do not provide effective downstream correction.
  • Statistical limitation: Mass responses are explicit but unreliable for population-level transport because their importance weights concentrate in sparsely populated stage-A tails, producing an extremely small effective sample.This separates functional interpretability from statistical usability in classifier-derived reweighting.
  • Scope and limitations: The study is limited to one leading-order dijet process, one setup per generator, eight high-level observables, generator-internal stage-A partons, and a density ratio conditional on a six-way common hard-event cohort.Pythia and Herwig are treated as fully specified configurations rather than tune-matched models or tests of individual microscopic parameters.

A Implementation details · A.1 Production configuration · A.2 Hard-event identity

The implementation uses a common MadGraph hard process with Pythia and Herwig configurations, while staged comparisons retain exact correspondence to the same underlying LHE events. Stage-specific generator logic and audited identity constructions define the production cohort.

  • A.1 Production configuration: The production sample combines MadGraph5_aMC@NLO 2.9.24, Pythia 8.312, Herwig 7.3.0, and the NNPDF23_lo_as_0130_qed PDF set.Pythia uses the Monash 2013 tune, whereas Herwig 7.3 uses its default parameter configuration.
  • A.1 Production configuration: The staged configurations separately switch showering, hadronization, and MPI or underlying-event modelling on or off.Stages A, B, and C correspond respectively to shower-only, shower plus hadronization, and shower plus hadronization plus MPI or underlying event.
  • A.2 Hard-event identity: The six generator representations are required to correspond to the same underlying LHE event, so HepMC event numbers are not used as the pairing key.This identity requirement is necessary for the staged analysis.
  • A.2 Hard-event identity: For Pythia, the production wrapper assigns an explicit hard-event identifier before showering and stores it in the HepMC event.The supplied passage continues with the identifier-assignment procedure, including ordinal handling across generator attempts.
  • A.2 Hard-event identity: The Pythia ordinal advances on every attempted next() call, including failed generator attempts, and successful and failed identifiers partition the attempted LHE sequence exactly.This makes failed attempts part of the audited identity sequence rather than silently skipping them.
  • A.2 Hard-event identity: For Herwig, hard-event identity is recovered from a canonicalized fingerprint of the complete MadGraph variation-weight vector, with no event-number, momentum-matching, or ordinal fallback.The Les Houches handler uses the signed variable-weight option to preserve this information, and both identity constructions are audited before forming the common cohort.

A.3 Processed data and functional export … B.2 Hard-event-clustered bootstrap

The analysis independently processes six generator-stage samples, exports validated additive KAN representations, and preserves hard-event pairing throughout cross-fitting, judging, and uncertainty estimation. Independent judge scores evaluate transported or recomposed weights without retraining, while clustered bootstrap resampling retains the paired generator evolutions of each hard scattering.

  • A.3 Processed data and functional export: Six generator-stage HepMC samples are processed independently, with eight observables and hard-event identifiers checked before constructing the common cohort.The samples comprise three stages for each of two event generators; development/test membership and fold assignment are then generated.
  • A.3 Processed data and functional export: Each exported additive KAN stores its bias, input standardization, source support, and tabulated centered response functions as a finite numerical object.The export is checked against the live KAN within its support and against the factorization in eq. (3.18).
  • B Statistical definitions and validation: The primary KAN models use cross-fitting, with the same hard-event grouping retained across all statistical resampling.Bootstrap replicas sample hard-event identifiers with replacement and include their associated generator representations.
  • B.1 Hard-event clustering and the independent judge: Closure and transport use a separate KAN judge cross-fitted on unweighted development samples at the target stage.The judge scores remain fixed, and transported or recomposed weights are applied only to Pythia events for weighted AUC evaluation.
  • B.1 Hard-event clustering and the independent judge: The judge is not retrained after reweighting, so weighted and unweighted comparisons use the same independent score function.This keeps evaluation separate from the reweighting procedure.
  • B.2 Hard-event-clustered bootstrap: Uncertainties are estimated by bootstrapping the underlying hard event because matched Pythia and Herwig representations originate from the same LHE matrix-element event.Representations at stages A, B, and C are likewise tied to the same hard scattering.
  • B.2 Hard-event-clustered bootstrap: Hard-event-clustered bootstrap resampling preserves each paired generator comparison, whereas row-wise resampling could underestimate statistical uncertainty.The clustered prescription measures fluctuations from varying the underlying sample of hard scatterings.

B.3 Effective sample size

The section evaluates transported-factor reliability using weight-based effective sample size and cumulative weight concentration. These diagnostics distinguish broad, statistically supported reweighting from apparent closure driven by rare, heavily weighted events.

  • Effective sample size: The effective sample size equals N for equal weights but can collapse when weights concentrate on a few events.This concentration reduces the statistical content of a large event sample to only a few effectively independent events.
  • Effective sample size: Transported factors are physically meaningful only when their apparent closure is supported across a sufficiently broad source-event population.Otherwise, rare events receiving very large importance weights can drive the observed agreement.
  • Effective sample size: The normalized ratio Neff/N measures statistical support: values near unity indicate broad reweighting, while very small values indicate a restricted phase-space tail.This adds reliability information that weighted AUC alone cannot provide.
  • Effective sample size: The cumulative normalized weight distribution reports how many highest-weight events are needed to accumulate 50% of the total weight.Hundreds or thousands of events suggest broad support, whereas one or two dominant events indicate a non-robust population-level transport.

C Functional Shapley implementation

The Shapley implementation evaluates the complete subset lattice for each additive-KAN fit, then aggregates feature-level allocations into observable sectors. Bootstrap intervals and efficiency checks quantify the numerical reliability of these allocations.

  • Subset and Shapley evaluation: All 256 subsets are evaluated by exact recomposition for each exported additive-KAN fit, with Shapley allocations computed from the complete subset lattice.The construction follows algorithm 2 and is repeated for each seed and value function.
  • Uncertainty estimation: Statistical intervals use 400 hard-event-clustered bootstrap replicas per seed.The intervals accompany the Shapley allocations obtained from the complete subset lattice.
  • Sector aggregation: Sector allocations are formed after computing Shapley values for the eight individual observables by summing the relevant feature-level values within each replica.The reported multiplicity, mass, and shape sectors therefore derive from feature-level allocations rather than being computed independently.
  • Numerical validation: The Shapley efficiency relation is checked for every seed and bootstrap realization, with the largest residual below 4×10−16.This provides a numerical consistency check on the allocation procedure.

D Weight support and influence diagnostics

Effective-sample and cumulative-weight diagnostics show that stage A mass factors cannot support reliable population-level transport measurements. Additional influence tests assess whether this reflects rare individual events or a persistent lack of statistical support.

  • Weight support diagnostics: Stage A mass factors lack sufficient support for reliable population-level transport measurements.This conclusion follows from effective-sample and cumulative-weight diagnostics.
  • Influence diagnostics: Additional influence tests distinguish rare-event effects from persistent statistical-support limitations.The tests are intended to determine which mechanism drives the observed behavior.

D.1 Tail support

Poor support in the stage-A mass tails is concentrated in a few rare events, unlike the broadly distributed multiplicity reweighting. Consistent training-seed results therefore do not establish adequate statistical support.

  • Mass tails: A leading-mass event near 39.0 GeV carries 49.1% of the total component weight, while the corresponding subleading-mass event carries 87.1%.The concentration is traced to sparsely populated stage-A mass tails rather than a broad source-distribution mismatch.
  • Multiplicity factors: For both leading and subleading multiplicity, the five largest weights together account for less than one percent of the corresponding total weight.Multiplicity factors are therefore qualitatively different from the concentrated mass-tail factors.
  • Multiplicity factors: Hundreds of development events are required to accumulate half of the reweighted multiplicity measure, so agreement across training seeds does not demonstrate statistical support.Several independently trained models can consistently identify the same sparsely populated structure.

D.2 Influence and targeted refitting tests

Influence and targeted-refitting tests show that mass-based responses are sensitive to influential events and fail statistical support tests, while multiplicity transport remains stable. The two mass factors fail differently: subleading mass has an event-sensitive tail, whereas leading mass remains concentrated across the high-mass tail.

  • Influence tests: 0.8970 versus 0.7642: the subleading-mass weighted AUC changes substantially when the dominant event is present versus absent, while leading multiplicity remains 0.5142 versus 0.5139.The mass diagnostics are sensitive to individual high-weight events, whereas supported multiplicity transport is substantially more stable.
  • Targeted refitting tests: Removing the dominant event changes the centered subleading-mass response range from approximately 19.33 to 17.52, exposing direct sensitivity in its learned tail.The targeted refit isolates the influence of that particular event on the subleading-mass response.
  • Targeted refitting tests: Removing the five most influential events changes the leading-mass response range only from approximately 22.36 to 22.52, indicating a more persistent concentration across the available high-mass tail.The leading-mass response is less altered by this targeted removal than the subleading-mass response.
  • Support limitations: Neither mass factor provides a statistically supported standalone transport weight, despite each yielding an explicit and reproducible response function.The subleading-mass response has a strongly event-sensitive tail, while the leading-mass response exhibits persistent high-mass-tail concentration.
Loading 2608.15952v1…