Source-linked AI summary

Uncovering Hidden Leptonic Correlations with Flow Matching and Autoencoders

Haruto Kitagawa, Satsuki Nishimura, Hajime Otsuka

arXiv:2608.15042v1hep-phcs.LGhep-th

TL;DR

Unresolved relationships among neutrino masses, mixing parameters, and CP phases motivate a global Type-I seesaw analysis using flow matching and autoencoders, which uncovers cluster-dependent nonlinear relations among leptonic observables.

  • Problem

    Relationships among neutrino masses, leptonic mixing parameters, and CP phases remain incompletely understood, alongside undetermined neutrino properties.

  • Method

    The study uses flow-matching posterior estimation to sample experimentally viable Type-I seesaw parameters and autoencoders to identify correlations among leptonic observables.

  • Results

    3 × 10^6 generated points yielded 640 solutions with χ2 < 45, while the autoencoder found four clusters and cluster-dependent relations involving CP phases and neutrino masses.

  • Takeaways & Limitations

    Combining flow matching with autoencoders reveals hidden nonlinear structures in high-dimensional leptonic distributions without imposing a specific flavor symmetry.

  • Takeaways & Limitations

    The analysis reports low regression accuracy for mixing angles, influenced by their narrow distributions and differing numerical treatment from masses and CP phases.

Abstract

from arXiv · show

We perform a global search for values of the Yukawa matrices and Majorana masses in the Type-I seesaw mechanism. Using flow matching, which is a generative artificial intelligence (generative AI) method, we generate a broad set of solutions reproducing the experimentally measured values of the neutrino mass-squared differences and the mixing angles. Then, a machine learning method known as an autoencoder is applied to uncover non-trivial correlations among physical quantities in the lepton sector. Our analysis reveals new non-linear relations involving neutrino masses and CP phases. These findings may contribute to elucidating the origins of the mass hierarchies and mixing patterns among generation structure.

1 Introduction

The study addresses unresolved questions in lepton flavor by globally exploring Type-I seesaw parameters compatible with low-energy neutrino data and identifying non-trivial physical correlations. It combines flow matching for parameter generation with an autoencoder for correlation discovery.

  • Motivation: The lepton sector’s neutrino mass scale, mass ordering, Majorana nature, and leptonic CP violation remain incompletely determined.These issues are part of the broader unresolved problem of explaining observed fermion mass and mixing hierarchies.
  • Contribution: Unlike approaches reconstructing particular high-energy parameters, this study explores the global distribution compatible with low-energy neutrino data and searches for emergent non-trivial correlations.The analysis focuses on the resulting parameter ensemble rather than one reconstructed parameter set.
  • Method: Flow matching enables global exploration of the high-dimensional Yukawa-matrix and Majorana-mass space, where allowed regions may be narrow, nonlinear, or disconnected.The method is motivated by the inefficiency of direct random scans when viable regions occupy only a small fraction of parameter space.
  • Method: An autoencoder compresses derived physical quantities into a lower-dimensional latent representation and uses reconstruction accuracy to identify correlations among them.Accurate reconstruction after dimensionality reduction indicates dependencies among quantities that would otherwise be lost if they were mutually independent.

2 Background

This section formulates Majorana neutrino masses through the Type-I seesaw mechanism and specifies the bases, mass ordering, and mixing conventions used in the analysis. It also summarizes experimental and cosmological constraints on lepton-sector observables and neutrino masses.

  • Seesaw framework: Majorana neutrino masses arise through the Type-I seesaw mechanism, with Y^ν and M denoting the neutrino Yukawa and Majorana mass matrices.The setup uses the charged-lepton flavor basis, where the charged-lepton Yukawa matrix is diagonal.
  • Mass basis: The numerical analysis uses a diagonal Majorana mass matrix, while the complex symmetric light-neutrino mass matrix is diagonalized by the unitary U_PMNS matrix.The light-neutrino eigenvalues m_i are chosen non-negative.
  • Mass ordering: The study focuses on normal ordering, defined by m_1 < m_2 < m_3, while the same analytical method also applies to inverted ordering.Inverted ordering is characterized by m_3 < m_1 < m_2.
  • Leptonic mixing: The leptonic mixing parameterization contains three mixing angles, one Dirac CP phase, and two Majorana phases.The Dirac phase is represented through the Jarlskog invariant, while Majorana-phase information is expressed through corresponding invariants.
  • Experimental constraints: NuFIT 6.0 provides the adopted experimental lepton-sector results, including 3σ C.L. intervals for the magnitudes of PMNS matrix elements.The fit incorporates Super-Kamiokande atmospheric-neutrino data.
  • External mass constraints: Cosmological and neutrinoless-double-beta-decay limits constrain the neutrino mass scale, including ⟨m_ee⟩ < 36 meV at 90% C.L. from KamLAND-Zen.DESI BAO plus CMB bounds depend on the assumed total-mass prior, and the effective Majorana mass is relevant for future nEXO sensitivity.

3 Parameter Optimization with Flow Matching

This section formulates seesaw-parameter determination as simulation-based inference for a non-unique inverse problem. Flow matching estimates p(G|Lexp) from sampled parameter–observable pairs, using a 23-dimensional parameter vector and 11-dimensional observable vector.

  • Problem formulation: The inverse map from seesaw parameters G to neutrino observables L is generally non-unique, so parameter determination is formulated as simulation-based inference.Samples from a prior over G are propagated through the Type-I seesaw relation to obtain observables.
  • Parameterization: The fitted parameter vector has dimension dG = 23, while the observable vector has dimension dL = 11.The 23 parameters comprise the real and imaginary parts of a general complex 3 × 3 Yukawa matrix, two heavy-mass ratios, a logarithmic scale parameter, and two phases; the observables comprise two mass-squared differences and nine absolute PMNS elements.
  • Conditioning choice: Conditioning on the nine absolute PMNS matrix elements produced more accurate flow-matching generation than conditioning directly on the three mixing angles in numerical tests.Although the nine quantities are not independent, they are used as conditioning variables for numerical accuracy.
  • Posterior estimation: The optimization targets the posterior distribution p(G|Lexp), whose samples provide parameters yielding observables close to the experimental values.This posterior-based strategy is intended to make the search for viable seesaw parameters efficient despite the inverse problem’s difficulty.
  • Flow-matching training: Training pairs (G1, L1) are generated by sampling many G1 instances from a uniform distribution and computing each L1 with the Type-I seesaw mechanism.The resulting parameter–observable pairs train the vector field used to represent the probability path.

1. Simulation

The simulation samples many G1 values uniformly and computes the corresponding observable L1 through the Type-I seesaw mechanism to construct training dataset D.

  • 1. Simulation: A large number of G1 values are sampled from a uniform distribution, and each corresponding L1 is calculated using the Type-I seesaw mechanism.These calculations produce the training dataset D.

2. Training

The training procedure learns a conditional vector field from training data by sampling conditional probability paths and minimizing a teaching-signal loss.

  • 2. Training: Training uses data D to learn the conditional vector field v_t,L(G; θ).
  • 2. Training: At each step, the procedure samples t ∼p(t) and draws G_t from the conditional probability path defined in Eq. (3.7).
  • 2. Training: The parameters θ are optimized to minimize the loss function in Eq. (3.6), using the vector field in Eq. (3.11) as the teaching signal.

3. Generation

The generation phase fixes the condition to the experimental value Lexp, samples an initial point from a standard normal distribution, and evolves it through a trained vector field. The resulting G1 is used as a sample from the posterior distribution p(G|Lexp).

  • Generation: Generation fixes the condition to the experimental value Lexp and samples G0 ∼N(0, I) from the standard normal distribution.The initial point is drawn before numerical evolution begins.
  • Generation: The method numerically solves the ordinary differential equation governed by the trained vector field vt,Lexp(G; θ) from t = 0 to t = 1.This evolution transforms the initial point into the generated sample G1.
  • Generation: The resulting G1 is used as a sample from the posterior distribution p(G|Lexp).The posterior is conditioned on the experimental value Lexp.

4. Evaluation

The evaluation targets representative ensembles because the inverse map is non-unique and may contain multimodal distributions and extended degeneracies. The Transformer flow generates many high-accuracy solutions, although accuracy varies across observables.

  • Evaluation objective: The generative model is evaluated as an ensemble generator because p(G|Lexp) can be multimodal and degenerate rather than selecting one best-fit parameter point.The forward map G 7→L is generally non-injective, so conditioning on Lexp does not uniquely determine G.
  • Model comparison: The Transformer architecture achieved higher generation accuracy than a simple fully connected network under otherwise comparable training settings.The improvement is attributed to self-attention modeling dependencies among different components.
  • Generated-sample accuracy: 956,607 generated samples satisfy χ2 < 2,500, including 640 samples satisfying χ2 < 45; the former set trains the autoencoder.The χ2 criterion evaluates agreement with the five observables simultaneously.
  • Generated-sample accuracy: 640 samples satisfy χ2 < 45 despite none of the 107 training data points meeting that criterion, showing that Flow Matching discovers solutions beyond observed training samples.The paper attributes this to capturing the high-dimensional and nonlinear geometry relating model parameters to physical observables.
  • Observable distributions: Generated marginal distributions cover broad, continuous regions, with ∆m2_31 near the preferred region and sin2 θ23 concentrated near its best-fit value, whereas ∆m2_21 is broader and asymmetric.A broad marginal does not necessarily imply global inaccuracy because χ2 constrains all five observables simultaneously.

4 Correlation Discovery via Autoencoder … 4.3 Numerical Implementation and Results

The paper uses a two-dimensional autoencoder latent space to discover correlations in seesaw observables generated from experimentally consistent parameters. Clustering reveals branch-dependent CP-phase structure, including δCP ≃ α31/2, while quadratic latent-space regressions approximate selected observables.

  • 4 Correlation Discovery via Autoencoder: Parameters satisfying χ2 < 2,500 are used to calculate physical observables through the Type-I seesaw mechanism.
  • 4 Correlation Discovery via Autoencoder: Because the generated data reproduce measured mass-squared differences and mixing angles, the paper applies autoencoders to extract non-trivial correlations from the nine-dimensional observables.
  • 4.1 Method: The autoencoder compresses input observables into a lower-dimensional latent space, where retained features can mediate relationships among physical quantities through regression.The latent dimension l is chosen below the input dimension d to avoid overfitting or identity mapping.
  • 4.1 Method: Several observables reconstruct accurately from a bottleneck, indicating concentration near a low-dimensional manifold rather than merely reflecting the imposed latent dimension.
  • 4.2 Network Architecture: The network uses a two-dimensional latent representation with fully connected layers, layer normalization, and four residual blocks, while the decoder approximately mirrors the encoder.The bottleneck enables compact correlation representation and direct latent-space visualization.
  • 4.2 Network Architecture: The choice l = 2 enables visualization in the (z1, z2) plane but does not establish an intrinsic dimension of two; similar clustering appears for l = 3.
  • 4.3.1 Training: 956,607 instances are split 90%/10% into training and validation data, with batch size 64, 50,000 training steps, and validation every 1,000 steps.The data are drawn from χ2 < 2,500 samples and preprocessed before training.
  • 4.3.2 Latent Space Analysis: HDBSCAN finds clustered high-density regions in latent space, whose strongest observable dependence is in CP phases, with δCP and α31 occupying branch-dependent intervals.For χ2 < 45 points, clusters separate phases with δCP ≃ α31/2; cluster 1 spans +1.43 – +3.14 rad and cluster 3 spans −3.13 – −0.642 rad.

Cluster 0 · Cluster 1

Figure 4 presents the distributions of observables across the clusters, with the supplied passage specifically labeled as part of Cluster 0.

  • Cluster 0: Figure 4 shows the distribution of observables in each cluster.The supplied passage is associated with the Cluster 0 subsection.

Cluster 2 · Cluster 3

The analysis identifies cluster-specific correlation surfaces in observable space, while showing that latent-coordinate coefficients are not physical invariants. It also distinguishes observables that separate clusters from those accurately reconstructed despite overlapping distributions.

  • Cluster 3: Physical predictions reside in invariant correlation surfaces among observables, not in the numerical values or parameterizations of latent coordinates.Invertible latent-coordinate changes alter the parameterization but preserve its image in observable space.
  • Cluster 3: m1 overlaps substantially across clusters but has a large R2, indicating continuous variation along a latent direction and precise latent-coordinate estimation.Cluster separation and regression accuracy therefore measure different characteristics.
  • Cluster 3: δCP and α31 distinguish clusters through distinct phase regions, while their within-cluster variations can remain smooth along latent coordinates.Clusters 0 and 2 show high R2 values for both phases; Clusters 1 and 3 have markedly different phase ranges.
  • Cluster 3: The three mixing angles have very low R2 across clusters because flow-matching and χ2 constraints restrict them to narrow ranges near experimental values.Their residual variability is weakly associated with the mass scales or CP phases defining cluster structure.
  • Cluster 3: A low R2 for a mixing angle does not necessarily imply a large absolute prediction error when its observed distribution is narrow.Masses and CP phases are normalized, whereas the mixing angles are used as inputs in the analysis.
  • Cluster 3: For samples satisfying χ2 < 45, parallelogram domains in latent space are mapped through quadratic functions to predicted observable surfaces compared with generated data.Figure 5 defines the domain approximations and Figure 6 displays the mapped correlations.
  • Cluster 3: Each cluster localizes permissible physical quantities on a two-dimensional correlation surface, providing simultaneous-observable consistency conditions beyond individual predicted ranges.These surfaces combine information from oscillation experiments, cosmological observations, and neutrinoless double beta decay searches.
  • Cluster 3: The learned correlation surfaces play a role analogous to flavor-symmetry sum rules but arise from an experimentally conditioned ensemble without imposing such rules.Table 4 reports the cluster-specific R2 coefficients for quadratic regressions in the latent variables.

5 Conclusions

The study combines flow-matching posterior estimation and autoencoders to explore Type-I seesaw parameters consistent with neutrino data and uncover hidden leptonic correlations. The analysis finds cluster-dependent nonlinear relations involving neutrino masses and CP phases, while identifying renormalization effects and sparse autoencoders as directions for future work.

  • Conclusions: Flow-matching posterior estimation generated Type-I seesaw parameter points constrained by experimentally measured neutrino mass-squared differences and mixing angles.The method learned the conditional distribution of free parameters without imposing a specific flavor symmetry on the neutrino Yukawa matrix.
  • Conclusions: 3 × 10^6 generated parameter points yielded 956,607 samples with χ2 < 2,500 and 640 samples with χ2 < 45.These figures quantify the generated sample set and the two selection criteria reported in the conclusions.
  • Conclusions: The autoencoder’s two-dimensional latent representation revealed four distinct clusters, with particularly pronounced separation in the CP phases.Cluster-dependent relations were obtained by approximating each cluster’s data distribution and expressing observables through latent variables.
  • Conclusions: The identified clusters imply distinguishable tendencies between δCP and both the neutrino-mass sum and the effective Majorana mass relevant to neutrinoless double-beta decay.These relationships arise from nonlinear correlations among the extracted leptonic observables.
  • Conclusions: Combining flow matching with autoencoders provides a data-driven framework for identifying experimentally consistent parameter points and extracting hidden structures from high-dimensional distributions.The conclusions characterize the reported relations as non-trivial findings produced by integrating data generation with feature analysis.
  • Conclusions: Future work should test whether the cluster structures and correlations remain stable under renormalization-scale evolution and assess sparse autoencoders for more interpretable latent features.The proposed sparse architecture uses an overcomplete latent representation in which only a small subset of latent variables is activated for each input.

A Formulation of flow matching

Flow matching learns a continuous-time probability-density flow by regressing on analytically known conditional velocities rather than explicitly computing the intractable marginal velocity field. Training is simulation-free, while ODE integration is deferred to inference to generate samples.

  • Continuous-time flow: Flow matching transports probability density from a tractable reference distribution p0 to a data distribution approximating q through an ODE-generated continuous-time flow.The velocity field is learned as a regression problem without numerical ODE integration during training.
  • Gaussian paths: Gaussian conditional paths connect a standard normal distribution at t = 0 to distributions concentrated near x1 at t = 1, with analytically determined conditional velocities usable as regression targets.For Gaussian paths, x0 is sampled from N(0, I), and the velocity follows by differentiating the conditional flow.
  • Conditional flow matching: Conditional flow matching avoids direct evaluation of the intractable marginal velocity by regressing on known conditional velocity fields, with equivalent training gradients up to a θ-independent constant.Each training step samples a data point, time, and conditional-path point, then performs squared-error regression without sequential ODE solving.
  • Conditional flow matching: Conditional paths act as virtual transport channels ending at data points, whose superposition over q produces the effective marginal velocity field.Conditional flow matching samples one component of this superposition in each mini-batch.
  • Generation: After training, the learned velocity field is integrated numerically from t = 0 to t = 1, making ODE solving an inference step rather than a training requirement.This separation avoids expensive trajectory integration during training while preserving a deterministic continuous-time generative process.
Loading 2608.15042v1…