Source-linked AI summary

A Spectral Identifiability Threshold for Dissipative Rate Recovery from Truncated Liouvillian Spectra

Yujun Ji, Somyajit Chakraborty

arXiv:2608.29302v1cs.LGphysics.comp-phquant-ph

TL;DR

The paper asks how much of a truncated, sorted Liouvillian spectrum is needed to identify dissipative rates when amplitude damping and dephasing have overlapping signatures. It analyzes a commuting Lindblad model with an affine closed-form spectrum and benchmarks recovery from slow non-steady modes. Population modes impose a lower bound k ≥D for uniform dephasing identifiability, with measured saturation at tested sizes but weaker advantages under perturbations and outside the noise-free simulator setting.

  • Problem

    The study addresses the unresolved spectral-sufficiency question of how much truncated, sorted Liouvillian-spectrum information is needed to recover dissipative rates.

  • Method

    The authors compute analytically structured spectra for a commuting Lindblad family, retain sorted slow modes, and compare linear and structure-agnostic tabular inverse estimators.

  • Results

    The dephasing threshold reaches the lower bound k⋆(γϕ) = D at n = 4, 5, 6, while tabular learners plateau around 10^-4 because tree ensembles cannot represent the affine map exactly.

  • Takeaways & Limitations

    The amount and structure of retained spectral information determine whether dissipative parameters are recoverable from sorted spectral prefixes within this commuting family.

  • Takeaways & Limitations

    The conclusions are established for noise-free simulator spectra of a known generator, while finite-shot measurement-derived spectra remain open.

Abstract

from arXiv · show

Open quantum systems lose energy and phase coherence through different dissipative processes, but these processes can produce overlapping dynamical signatures. The Liouvillian spectrum summarizes how such a system relaxes, yet it is not obvious how much of that spectrum is needed to distinguish the underlying dissipation rates. We study this question for amplitude damping and dephasing in a six-qubit Lindblad model whose spectrum can be derived analytically. We retain only the slowest non-steady spectral modes and ask how many are required before each dissipative rate becomes recoverable. We show that population modes contain no dephasing information, which creates a lower bound of D = 2^n retained modes for uniform dephasing identifiability in the relevant rate regime. The measured recovery threshold reaches this bound at n = 4,5,6, while n = 3 remains above it. At n = 6, least squares achieves a mean joint absolute error of order 10^-9, compared with 4.355 x 10^-4 for four tabular learning methods. Robustness tests show that this advantage weakens when the spectra are perturbed and when a transverse field breaks the commuting structure. These results show that the amount and structure of retained spectral information can determine whether dissipative parameters are recoverable, independently of the estimator used. The present conclusions apply to noise-free simulator spectra rather than measurement-derived spectra.

I. INTRODUCTION

The paper asks how much of a sorted, truncated slow Liouvillian spectrum is needed to recover amplitude-damping and dephasing rates. It combines an analytically solvable commuting model with inverse benchmarks, showing that spectral structure and retained depth govern identifiability and estimator performance.

  • I. INTRODUCTION: The unresolved question is how much of a simulator-supplied truncated Liouvillian spectrum must be retained before dissipative rates become identifiable.The focus is spectral sufficiency rather than recovery from measurement records or experimentally collected data.
  • I. INTRODUCTION: The benchmark assigns amplitude-damping and dephasing rates, computes the Liouvillian spectrum, removes the steady mode, and uses the slowest k real-valued modes as inverse-task inputs.The target is recovery of both coefficients from sorted, mode-unlabelled spectral features.
  • A. A closed-form spectrum and where each rate becomes identifiable: Every eigenvalue is affine in the two dissipative rates, and exact numerical checks support the closed-form spectrum and its structural near-linear identifiability.At n = 4, the maximum entrywise deviation was 1.67×10^-16 with an exactly vanishing upper-triangular residual.
  • A. A closed-form spectrum and where each rate becomes identifiable: Between k = 60 and k = 70, the mean joint MAE drops from 1.32 × 10^-3 to 8.56 × 10^-10 as dephasing information enters the retained spectrum.Least squares reaches a ≈10^-9 quantization-limited regime at larger budgets.
  • A. A closed-form spectrum and where each rate becomes identifiable: k⋆= 13 for γ1 and k⋆= 64 for γϕ are the smallest held-out-MAE thresholds below 10^-7, stable across split seeds 42, 43, and 44.The finer scan resolves the dephasing onset at k⋆(γϕ) = 64 = D, while seed agreement checks partition stability rather than independent replication.
  • A. A closed-form spectrum and where each rate becomes identifiable: Ridge reproduces the same retained-sector behavior, while underdetermined high-budget fits require interpreting its values together with the selected penalty.This makes the spectral-sector effect distinct from a single least-squares implementation choice.

B. A derived lower bound on the retained window

The commuting model's population sector is blind to dephasing, so uniformly recovering γϕ from sorted prefixes requires retaining at least D = 2^n modes. This lower bound is a non-injectivity result for worst-case rate pairs, while attainment remains an empirical question.

  • B. A derived lower bound on the retained window: Population modes have Re λ = −γ1n_i and do not depend on γϕ; after removing the steady mode, D −1 non-steady population modes remain dephasing-blind.When populations precede coherences, these modes occupy the leading sorted prefix.
  • B. A derived lower bound on the retained window: The first γϕ-informative coordinate varies from 1 to D across the sampled rate pairs because the population-first ordering depends on γϕ/γ1.At n = 6, 120 of 1000 sampled pairs attain the worst-case position D.
  • B. A derived lower bound on the retained window: k ≥D is required for uniform dephasing identifiability because two distinct γϕ values at fixed γ1 produce identical first D −1 sorted coordinates in the population-first regime.The bound applies to any estimator using sorted prefixes, but it guarantees only that dephasing information is present, not that the resulting system is solvable or well-conditioned.
  • B. A derived lower bound on the retained window: The γ1 onset is also rate-order dependent: the leading coordinate equals −γ1 on most samples, but coherences become slower when γϕ < γ1/4.This explains why a one-coordinate estimator is not exact and why γ1 requires separating slow-mode identities.

C. System size and the role of commutation

The dephasing-identifiability threshold reaches the structural bound D = 2^n for n = 4, 5, and 6, while transverse-field perturbations eventually disrupt the commuting-model behavior. Structure-agnostic tabular learners plateau far above least squares, and additional spectral coordinates provide little benefit under this protocol.

  • System-size threshold: The dephasing threshold equals D = 2^n at n = 4, 5, and 6, whereas n = 3 requires k⋆ = 10 above D = 8.The required fraction of the non-steady spectrum shrinks from 15.9% to 1.6% between n = 3 and n = 6.
  • Commutation robustness: At n = 4, the transverse-field threshold k⋆(γϕ) = D = 16 survives through h = 10^-2, but neither threshold is reached within k ≤ 32 at h = 10^-1.The h = 10^-2 γϕ result is marginal because its held-out MAE at k = D is 8.6 × 10^-8, close to the 10^-7 cutoff.
  • Structure-agnostic learners: The benchmarked tabular learners plateau at mean joint MAE values from 4.355 × 10^-4 for AutoGluon to 6.086 × 10^-4 for CatBoost.The four families agree within 24%, while their ordering is not treated as a definitive library ranking.
  • Estimator structure: The learner gap reflects a hypothesis-class mismatch: tree ensembles are piecewise constant, whereas a linear estimator matches the affine rate map exactly.This comparison is a protocol-specific reference point rather than a claim about the ceiling of automated machine learning.
  • Spectral budget: All four tabular families achieve their best aggregated joint MAE at k = 80, while extending the retained spectrum to k = 400 increases joint MAE by 58.6%.For the present distribution and protocol, the slowest 80 non-steady modes already carry the useful rate information.
  • Evaluation caveats: The learner comparison is descriptive because nested budgets, overlapping splits, and heavy-tailed residuals limit the effective sample size and complicate model ordering.The effective sample size is closer to five budgets than fifteen blocks, and residuals at the AutoGluon k = 80 operating point are heavy-tailed.

E. What the sorted coordinates carry

Sorted Liouvillian coordinates are useful predictive features, but degeneracies and correlations limit their interpretation as physical mode identities. The nearly linear inverse structure explains why simple linear estimators can outperform nonlinear tabular learners, while noise and model-specific diagnostics constrain the conclusions.

  • Spectral-coordinate interpretation: Real spectral coordinates recover the targets almost as well as the full feature set, whereas imaginary coordinates alone are weak.For fixed physical modes, imaginary parts are rate-independent, but sorting transfers indirect rate information through rate-dependent ordering.
  • Robustness: Noise reduces the linear advantage from 56.2 at η = 0 to 1.19 at η = 10^-1, making the advantage specific to near-exact spectra.The ridge-versus-random-forest scan bounds persistence under perturbation rather than reproducing the larger least-squares separation.
  • Estimator implications: The inverse map is close to linearly identifiable, so least squares and ridge strongly outperform the saved AutoML and boosted-tree runs once enough modes are retained.The four-family benchmark therefore measures nonlinear learner behavior on a nearly invertible direct-spectrum task.
  • Interpretability limits: Feature attributions should not be treated as isolated physical mechanisms because sorted coordinates are correlated, degenerate, and can change rank near crossings.The SHAP analysis audits one fitted XGBoost model rather than establishing the central identifiability claim.
  • Spectral-coordinate interpretation: Figure 3 identifies near-degenerate sorted coordinates as a rank-stability warning rather than globally tracked physical eigenmodes.Additional eigenvector tracking would be required to assign stable physical mode identities.

III. DISCUSSION

The discussion frames the result as a spectral-sufficiency finding for a controlled, commuting Liouvillian inverse problem. The identifiability threshold is sharp in the tested regime, but robustness and experimental scope remain limited.

  • The study concerns direct Liouvillian-spectrum inversion rather than experimental recovery from finite measurement records.
  • The measured threshold k⋆(γϕ) = D is attained for n = 4, 5, 6 in the commuting family, while the result is conditional on the sampled rate ensemble.
  • The derived requirement k ≥ D follows because population modes are blind to γϕ and sorted prefixes shorter than D are non-injective in γϕ.
  • The claims are limited to synthetic, noise-free spectra from a fixed commuting model and do not establish a device-ready rate-estimation protocol.
  • 4.1 × 10^-8 and 4.3 × 10^-8 are ridge MAEs on high-rate and diagonal holdouts, whereas the boundary-margin holdout reaches 5.8 × 10^-4.

IV. METHODS

The methods section’s supplied passage documents artifact reproducibility through archived project files and an automated claims audit.

  • All cited artifact paths resolve within the archived project release, whose automated audit re-checks reported metrics, thresholds, and summary statistics.
  • The audit table records each checked claim together with its source file.
  • The supplied methods evidence focuses on release-level verification rather than the model equations or spectral construction.

A. Markovian open-system dynamics

The model assumes time-homogeneous Lindblad dynamics for a six-qubit system with local amplitude damping and dephasing. Its commuting, computational-basis-diagonal structure is central to the reported identifiability behavior.

  • The density operator evolves under time-homogeneous, completely positive, trace-preserving semigroup dynamics in an n-qubit Hilbert space of dimension D = 2^n.
  • For the studied system, n = 6, ω = 1.0, and J = 0.05, with units scaled by the single-qubit level splitting.
  • Local amplitude damping and dephasing act through collapse operators with coefficients γ1 and γϕ on each qubit.
  • For an isolated qubit, dephasing contributes dρ01/dt = −2γϕρ01, while the combined coherence decay rate is γ1/2 + 2γϕ.
  • The Hamiltonian and dephasing operators are diagonal in the computational basis, placing the benchmark in the commuting, purely σz limit.

B. Liouvillian spectrum

The Liouvillian spectrum is computed from a thresholded, sorted set of non-steady eigenvalues and represented by real and imaginary coordinates. Triangularity in the computational-pair basis makes the eigenvalues analytically accessible.

  • The eigenvalue problem is L v_m = λ_m v_m with λ_m ∈ C and v_m ∈ C^(D^2).
  • Modes with Re(λ_m) < −10^-10 are retained, excluding the steady mode and leaving D^2 − 1 non-steady modes.
  • Retained eigenvalues are ordered by decreasing real part and then ascending signed imaginary part, using the latter only as a deterministic tie-breaker.
  • The sorted spectral input is represented as x_k = (a_1, …, a_k, b_1, …, b_k)^T ∈ R^(2k).
  • Ordering computational-pair basis elements by excitation content makes the generator triangular, so its eigenvalues are its diagonal entries.
  • The closed form was verified entrywise at n = 4 with maximum absolute deviation 1.67 × 10^-16 and zero upper-triangular residual.

D. Inverse regression task

The inverse task uses truncated Liouvillian eigenvalue features to recover amplitude-damping and dephasing coefficients from controlled synthetic datasets. Predictors are selected spectral coordinates, with input dimension p = 2k and matched feature budgets up to k = 400.

  • Dataset and targets: Each sample contains independently sampled γ1 and γϕ coefficients, with one regressor trained for each target.Both coefficients are sampled uniformly from [0.005, 0.05].
  • Spectral features: The synthetic datasets are generated by constructing the Hamiltonian and collapse operators, forming the Liouvillian, computing its eigenvalues, and applying the prescribed sorting rule.The Hamiltonian is constructed once and reused across sampled rate pairs.
  • Dataset and targets: The matched comparison uses five files containing 1000 unique rate pairs at fixed n = 6, ω = 1.0, and J = 0.05.Both coefficients lie in [0.005, 0.05], and the ordered rate-pair sequence is reproduced with seed 0.
  • Spectral features: Only eig_real_ and eig_imag_ columns are used as predictors, while provenance fields k, n, omega, and J are excluded.The target columns are gamma1 and gammaphi.
  • Feature budgets: The matched cross-model feature budgets are k ∈ {80, 160, 240, 320, 400}, corresponding to input dimensions p ∈ {160, 320, 480, 640, 800}.A separate AutoGluon study examines k ∈ {10, 20, 30, 40, 50, 60, 70} as a lower-resolution scan.

G. Training, validation, and test protocol

Training and evaluation use deterministic, repeated partitions of the same 1000 synthetic rate pairs, with separate handling for benchmark models and the ordinary-least-squares threshold scan. Diagnostics use a fixed k = 80 dataset and the same split seeds.

  • Data partitioning: Each benchmark split contains 600 training, 200 validation, and 200 test samples after two deterministic partitions of 1000 samples.The procedure is repeated for split seeds 42, 43, and 44.
  • Evaluation protocol: Model selection and early stopping use only validation data, while final metrics are computed on held-out test partitions.The test partition does not enter fitting or hyperparameter search.
  • Model families: The matched benchmark compares XGBoost, LightGBM, CatBoost, and AutoGluon, fitting separate regressors for γ1 and γϕ before joint scoring.The first three are gradient-boosted decision-tree methods, whereas AutoGluon assembles and tunes candidate predictors.
  • Diagnostics: The diagnostic suite uses the k = 80 dataset and split seeds 42, 43, and 44 for spectral, perturbation, holdout, and error-map analyses.The archived analysis code fixes the coordinate selections, noise model, seeding, and holdout masks.
  • Scaling and threshold scan: The scaling study regenerates datasets at n = 3, 4, and 5 and applies the same real-coordinate OLS threshold scan used at n = 6.The threshold k⋆ is the smallest budget with held-out MAE below 10^-7 across the specified split seeds.
  • Robustness: At n = 4, the transverse-field robustness scan varies h across six values and records thresholds above k = 32 as right-censored rather than extrapolating them.At h = 0, the control thresholds are k⋆(γ1) = 9 and k⋆(γϕ) = 16.

K. Hyperparameter optimization

Hyperparameter selection and evaluation use target-wise validation errors, explicit joint metrics, and matched paired comparisons across feature budgets and split conditions. Statistical conclusions are treated as exploratory because the repeated conditions are dependent.

  • Optimization: The three gradient-boosting methods use 60 Hyperopt/TPE evaluations per run over learning-rate, tree-depth, regularization, and sampling parameters.AutoGluon receives the same 60-evaluation budget using four preset-and-seed settings.
  • Metrics: Target-wise MAE, RMSE, and coefficient of determination are computed over N = 200 held-out test samples for each target.The targets are indexed by q ∈ Q, with predictions and true coefficients defined for every sample.
  • Metrics: Joint MAE is the arithmetic mean of the two target-wise MAEs.The two targets receive equal weight because they share units and sampling intervals.
  • Metrics: Joint RMSE pools all 2N target-sample errors before applying the square root, rather than averaging target-wise RMSE values.The square root is applied only after pooling squared errors.
  • Statistical comparisons: Model comparisons use 15 matched feature-budget and split conditions with paired tests selected by Shapiro–Wilk normality checks.Six pairwise comparisons receive Holm and Benjamini–Hochberg corrections, with Cohen’s dz and Cliff’s delta as effect sizes.
  • Statistical comparisons: The statistical tests and intervals are exploratory because feature budgets are nested and repeated splits overlap on the same rate pairs.Confirmatory inference would require independently generated datasets.

N. Interpretability analysis

Interpretability is a post-hoc analysis of saved XGBoost models at k = 80, combining feature-attribution, permutation, and correlation diagnostics. The archived implementation records model and environment details, while AI tools were used under author direction for coding and editing rather than for reported quantities.

  • Interpretability scope: Interpretability uses Tree-SHAP, ten-repeat permutation importance, and clustered Spearman correlations on XGBoost k = 80 split-1 models.The analyses use the 200-row test partition and are not treated as causal evidence.
  • Interpretability scope: Because spectra are sorted independently for each sample, spectral coordinates are operational feature indices rather than tracked physical eigenmodes.This constrains how feature-level interpretations should be read.
  • Reproducibility: The archived matched runs record versions for XGBoost, LightGBM, CatBoost, and AutoGluon, with additional environment details in ENVIRONMENT.md.The runners use CUDA mode, although individual AutoGluon constituents may execute on CPU.
  • Reproducibility: All custom simulation, dataset, modeling, statistical, and figure-generation code is archived on Zenodo with an automated claims audit.The release is identified by DOI 10.5281/zenodo.22160077.
  • Research process: AI systems were used under author direction for manuscript editing, simulation and analysis code, LaTeX formatting, and consistency checks, not to generate or alter reported quantities.The reported numbers were produced by the archived code.
Loading 2608.29302v1…