Source-linked AI summary
When Similarity Is Interaction-Driven: Quantum Kernels for Regime-Sensitive Learning
Hanqiu Peng, Jianlong Lu, Ying Chen
TL;DR
Fraud and anomaly-detection similarity can be governed by sparse interactions rather than distance alone, creating a need for interaction-aligned representations. The paper introduces a thin-slab model and entangled Pauli-string quantum kernel that encode high-order block interactions, then evaluates them synthetically and on fraud benchmarks. The kernel consistently outperforms conventional and engineered-interaction baselines in the evaluated regimes, while its classically evaluable construction demonstrates representational rather than computational quantum value.
Problem
Fraud and anomaly-detection labels can change across interaction-sensitive boundaries even when observations remain close in the original feature space.
Method
The paper combines a thin-slab interaction framework with an entangled Pauli-string feature map that encodes sparse high-order block products into an exactly block-factorized fidelity kernel.
Results
Across third-, fourth-, sixth-, and eighth-order synthetic tasks, the kernel consistently outperforms standard kernels and InterFea; it achieves strongest F1 on Credit Card Fraud Detection and ranks second on IEEE-CIS.
Takeaways & Limitations
Quantum-kernel performance depends on alignment between feature-map geometry and predictive structure, with the demonstrated gains being representational and statistical rather than computational quantum speedup.
Takeaways & Limitations
The controlled synthetic mechanism is deliberately aligned with the proposed feature map, while real-data interactions are unknown and IEEE-CIS does not show uniform dominance.
Abstract
from arXiv · showhide
Similarity in many decision systems is governed not by distance alone but by interactions among variables. In fraud and anomaly detection, small local perturbations can cross interaction-sensitive decision boundaries while leaving ambient distance almost unchanged. Motivated by this setting, we introduce a thin-slab interaction model and an interaction-driven quantum kernel constructed from entangled Pauli-string feature maps. The feature map explicitly encodes sparse high-order block interactions. We show that the resulting fidelity kernel is positive semidefinite, admits an exact block-factorized formulation, and induces a geometry sensitive to changes in interaction regime. Across balanced and imbalanced synthetic experiments spanning third-, fourth-, sixth-, and eighth-order interactions, the proposed kernel consistently outperforms linear, radial basis function, Laplacian, and polynomial kernels, as well as an engineered-interaction linear baseline supplied with the planted block products. On real fraud-detection benchmarks, it achieves the highest mean accuracy and F1 on Credit Card Fraud Detection and ranks second on IEEE-CIS Fraud Detection. These findings show that quantum-kernel performance depends on alignment between feature-map geometry and the underlying predictive structure, rather than on Hilbert-space dimension alone. Because the prescribed block-factorized kernel can also be evaluated exactly on a classical computer, the results establish predictive and representational value rather than computational quantum speedup.
1 Introduction
The paper frames fraud and anomaly detection as interaction-sensitive learning problems and introduces a quantum kernel whose geometry is aligned with sparse high-order interactions. Across synthetic and real evaluations, the method shows predictive gains while establishing representational rather than computational quantum advantage.
- Motivation: Fraud similarity can depend on sparse interactions among contextual variables, so nearby observations may occupy different risk regimes despite similar original features.This motivates similarity measures that respond to interaction structure rather than coordinate-wise proximity alone.
- Method: The proposed entangled Pauli-string feature map encodes fixed sparse high-order block products and yields a positive-semidefinite, exactly block-factorized fidelity kernel.The construction targets feature-map alignment rather than Hilbert-space dimension alone and can be evaluated efficiently under the known block partition.
- Synthetic evaluation: Across third-, fourth-, sixth-, and eighth-order synthetic tasks, the quantum kernel consistently outperforms standard kernels and the engineered-interaction linear baseline.In a balanced setting accuracy rises from approximately 0.50–0.52 to above 0.75; in imbalanced settings F1 rises from approximately 0.20–0.28 to above 0.53.
- Real-data evaluation: On real fraud benchmarks, the method achieves the strongest F1 on Credit Card Fraud Detection and remains competitive on IEEE-CIS Fraud Detection.Credit Card Fraud Detection mean F1 improves from 0.348 for the best competing classical kernel to 0.472, while IEEE-CIS ranks second to the Laplacian kernel.
- Problem setting: The thin-slab interaction model separates interaction-driven decision boundaries from generic distance-based similarity.It provides a controlled abstraction in which Euclidean proximity is misaligned with predictive similarity.
- Conclusion: The study concludes that quantum-kernel performance depends on alignment between feature-map geometry and predictive structure, not Hilbert-space dimension alone.Because the prescribed block-factorized kernel is classically evaluable, the demonstrated benefit is predictive and representational rather than computational quantum speedup.
2 Related Work
Related work motivates interaction-aligned similarity by showing that generic quantum embeddings and classical distance geometries can miss sparse, high-order predictive structure. The paper addresses this gap by encoding block-wise interactions directly into quantum-induced similarity geometry for fraud-oriented settings.
- Quantum kernels: Quantum-kernel research increasingly emphasizes feature-map design, data structure, and induced geometry over generic representational richness.This literature includes work on generalization, effective dimension, and concentration phenomena.
- Classical interaction learning: Classical interaction-learning methods decompose, select, or detect higher-order structure, but sparse interaction discovery becomes difficult as candidate combinations grow.Functional ANOVA, sparse hierarchical models, and tree-based methods represent complementary approaches.
- Positioning: The proposed quantum kernel differs by encoding structured block-wise interactions directly into entangled feature-map phases instead of explicitly enumerating interaction features.Its similarity therefore depends on hidden interaction regimes rather than Euclidean proximity in the original feature space.
- Application setting: Fraud detection is a natural application because subtle behavioral combinations can determine predictive regimes while fraudulent transactions remain locally similar to legitimate activity.Relevant contextual structure can include device identity, timing, behavioral history, and transaction sequence.
- Benchmarking: The study uses Credit Card Fraud Detection and IEEE-CIS Fraud Detection from the Fraud Dataset Benchmark for standardized evaluation beyond synthetic data.The real-data experiments are not intended to claim that fraud labels exactly follow the synthetic thin-slab mechanism.
3 High-Order Interaction Learning under Thin-Slab Geometry
The controlled problem tests interaction-aware learning when labels depend on planted block products and thin-slab sampling makes Euclidean proximity misrepresent regime similarity. It compares standard kernels and an engineered-interaction baseline against this interaction-driven structure.
- 3.1 Block-structured high-order interaction signal: Labels are generated from weighted majority votes over block-level regime bits defined by k-way products of coordinates.The model partitions d features into B disjoint blocks of size k and assigns each block a regime bit from its product sign.
- 3.1 Block-structured high-order interaction signal: For k > 2, the target signal is absent from individual coordinates and is not directly represented by polynomial kernels with degree below k.The relevant information is encoded in planted k-way monomials rather than first-order coordinates or pairwise distances.
- 3.2 Thin-slab distribution: Thin-slab sampling places most coordinates near zero, concentrating observations near coordinate hyperplanes instead of in a full-dimensional uniform cloud.The support is a union of thin slabs, with inactive coordinates of order ϵ.
- 3.2 Thin-slab distribution: O(ϵ) perturbations can flip a block regime, so nearby samples may have different labels despite nearly unchanged Euclidean distance.A sign change in one inactive coordinate reverses the block product while changing the Euclidean norm by only O(ϵ), creating a mismatch between distance and interaction geometry.
- 3.3 Why generic kernels can be misaligned: Distance-based kernels can smooth across interaction boundaries, while generic polynomial kernels embed the planted terms in a much larger unstructured expansion.For d = 20 and p = 4, the polynomial space contains 10,626 monomials up to degree four, while only five are planted block products.
- 3.3 Why generic kernels can be misaligned: InterFea augments raw features with planted block products and fits a regularized linear classifier directly in that engineered feature space.This provides a baseline with explicit access to the planted interactions without explicitly constructing its Gram matrix.
4 Quantum Feature Map and Kernel Construction
The feature map encodes sparse high-order block interactions through entangled Pauli-string phases, yielding a fidelity kernel whose geometry responds to block-product changes and whose structure factorizes across blocks.
- Feature-map design: The feature map compares samples through block products, so interaction-variable differences directly alter quantum phases and the resulting fidelity kernel.Its advantage is predicted only when the encoded block-product structure matches the data-generating mechanism.
- Feature-map design: Each block uses one qubit per coordinate, Hadamard preparation, one-body phase terms, and a k-way Pauli-string phase encoding the block interaction.All phase gates commute because they are diagonal in the computational basis.
- Kernel properties: Block factorization permits separate evaluation of the B block states and exact classical computation of the prescribed kernel.The same factorization also reduces simulated peak width from d to k wires.
- Kernel properties: The fidelity kernel is positive semidefinite and depends on both coordinate-wise differences and block-product discrepancies.The block contribution has an explicit closed-form dependence on interaction-regime changes.
- Circuit resources: For B = 10, increasing interaction order from k = 6 to k = 8 raises reusable-qubit width from 6 to 8 and logical CNOTs from 100 to 140 per observation.The engineered classical representation simultaneously grows from 70 to 90 coordinates.
- Circuit resources: Statevector storage scales as O(B2^k) per observation, while a cached pairwise entry costs O(B2^k) or O(Bk) using the closed form.These resource counts describe the represented logical circuit, not hardware-executed gates.
5 Experimental Design
The experiments test interaction-aligned kernels on controlled synthetic regimes and imbalanced fraud benchmarks, using matched validation protocols, noisy-training conditions, and engineered-interaction baselines.
- Synthetic settings: Synthetic experiments vary slab thickness ϵ ∈ {0.05, 0.15} and training-label flip probability pflip ∈ {0.00, 0.10}, while test labels remain clean.This isolates robustness to noisy supervision from contamination of the evaluation target.
- Synthetic settings: Balanced synthetic labels use threshold τ = 0 with a theoretically 50% clean positive-class proportion.The construction relies on symmetric block-regime bits and an odd number of blocks.
- Evaluation protocol: Each synthetic configuration uses 2000 training observations, 400 validation observations, and an independent 500-observation test set across ten random seeds.Balanced runs are selected and reported by accuracy, whereas imbalanced runs use class-balanced weights and validation/test F1.
- Higher-order scaling: The higher-order study fixes B = 10 while increasing interaction order, using two imbalanced settings with ϵ = 0.15, pflip = 0, and τ = 3.This isolates interaction-order effects from changes in the number of planted blocks.
- Baselines: Baselines include linear, RBF, Laplacian, degree-2/4/6/8 polynomial kernels, and InterFea, which receives the planted block products directly.All methods tune their hyperparameters under the common validation protocol before refitting on the full training sample.
- Real-data benchmarks: Real-data evaluation uses Credit Card Fraud Detection and IEEE-CIS Fraud Detection, both highly imbalanced binary classification benchmarks.Accuracy and F1 are reported, with validation F1 as the tuning objective and split-first preprocessing.
6 Results
Across balanced, imbalanced, noisy, higher-order, and real-data evaluations, the quantum kernel generally achieves the strongest performance, with gains over generic kernels and InterFea varying by setting.
- Balanced synthetic setting: The quantum kernel achieves the best mean accuracy in all four balanced configurations for both the smaller and larger synthetic settings.In the larger (d, k, B) = (20, 4, 5) setting, its gains over InterFea are modest but consistent.
- Learning-curve analysis: 0.718 to 0.753: quantum-kernel mean test accuracy rises across N = 250 to N = 2000 while remaining highest at every investigated sample size.InterFea improves from 0.666 to 0.734, whereas standard kernels remain near the 0.5 random-classification baseline.
- Imbalanced synthetic setting: The quantum kernel achieves the best mean F1 in all imbalanced synthetic configurations, with InterFea as the closest competitor.F1 is the primary metric because these experiments are imbalanced.
- Scaling to higher interaction order and dimension: At k = 6 and k = 8, the quantum kernel ranks first in both accuracy and F1, achieving (0.706, 0.412) and (0.695, 0.395), respectively.Relative to InterFea, the paired mean F1 gaps are 0.032 and 0.075.
- Scaling to higher interaction order and dimension: The quantum kernel outperforms InterFea across k ∈{3, 4, 6, 8}, with mean F1 gaps of approximately 0.054, 0.029, 0.032, and 0.075.The gap is positive at every evaluated interaction order and largest at k = 8.
7 Discussion
The discussion attributes the synthetic advantage to interaction-aligned fidelity geometry that persists with higher order and dimension, while real-data results and computational considerations delimit the claim.
- Interaction-aligned predictive value: Across third-, fourth-, sixth-, and eighth-order synthetic tasks, the Pauli-string fidelity kernel outperforms generic kernels and InterFea.The product-of-fidelities geometry adds predictive value beyond a linear representation of the planted variables.
- Quantum-specific contribution: The quantum-specific contribution combines block-factorized Pauli-string encoding of planted products with nonlinear fidelity composition across blocks.Its advantage over InterFea indicates that the improvement is not exhausted by appending planted products to a classical linear feature vector.
- Scope of the claim: The scaling study provides predictive and geometric evidence rather than computational quantum advantage because the prescribed block-factorized kernel has an exact efficient classical evaluation.A computational advantage would require a non-factorizing or otherwise classically hard feature map and an explicit complexity separation.
- Real-data nuance: On Credit Card Fraud the quantum kernel leads, but on IEEE-CIS the Laplacian kernel leads and the quantum kernel ranks second.These results indicate that usefulness depends on resemblance between real predictive structure and the circuit’s encoded interaction geometry.
- Limitations: The controlled synthetic study deliberately aligns its data-generating mechanism with the proposed feature map, limiting how directly its advantage transfers to unknown real interactions and block structures.IEEE-CIS is not uniformly dominated by the quantum kernel.
- Scalability and future work: The construction scales to larger d and k without enumerating all degree-k interactions, but dense kernel storage remains quadratic in sample size.Exact kernel-entry evaluation costs O(Bk) under the block-factorized formulation.
8 Conclusion
The study links quantum-kernel performance to alignment between feature-map geometry and sparse interaction structure, with consistent gains across synthetic interaction orders and strong fraud-benchmark results. It also emphasizes predictive and representational value rather than computational quantum advantage.
- The block-factorized quantum feature map encodes sparse block-wise k-way products as Pauli-string phases, yielding an explicit interaction-aware fidelity geometry.
- Across third-, fourth-, sixth-, and eighth-order settings, the proposed kernel consistently outperforms generic classical kernels and InterFea.Its advantage over InterFea is positive at every evaluated order and largest at eighth order.
- On fraud benchmarks, the method achieves the strongest F1 on Credit Card Fraud and ranks second on IEEE-CIS.
- The results support a quantum-induced predictive advantage in the studied interaction-driven regimes, while the factorized kernel remains exactly classically evaluable.The findings therefore establish predictive and representational value, not computational quantum speedup or universal superiority.
- Synthetic data are generated programmatically, and the real benchmark datasets are publicly available from their respective sources and through the Fraud Dataset Benchmark.
- Code is available from the authors upon reasonable request, and all authors contributed to study design, analysis, and manuscript writing.
A Numerical kernel implementation
The appendix distinguishes the main-text feature-map construction from additional numerical implementation details needed to reproduce the reported computations.
- The appendix records numerical details connecting the mathematical feature map, closed-form block overlap, circuit resources, and experimental settings to implementation.
A.1 Numerical backends
Most experiments generate block states with PennyLane’s default.qubit simulator, while the 80-dimensional setting uses exact NumPy amplitudes without changing the kernel definition.
- PennyLane’s default.qubit simulator generates block states for the d = 9, d = 20, real-data, and (d, k, B) = (60, 6, 10) experiments.
- For (d, k, B) = (80, 8, 10), exact NumPy amplitudes replace repeated simulator decomposition of a broadcast eight-qubit Pauli rotation.
- The NumPy evaluation path leaves the feature map and kernel definition unchanged and agrees with PennyLane outputs to below 10−10 maximum absolute error.
A.2 Batched Gram-matrix construction
The implementation constructs the full kernel by reshaping inputs into blocks, computing block overlaps in batches, and combining them with Hadamard products while enforcing unit self-fidelity.
- Inputs X ∈Rn×d are reshaped into B = d/k blocks, whose block-b statevectors are stored as rows of Sb(X) ∈Cn×2k.
- For datasets X and Z, each block Gram matrix is evaluated by batched matrix multiplication and elementwise squared modulus.
- The full kernel matrix is accumulated using the Hadamard product across block contributions.
- The training Gram diagonal is set to one for exact unit self-fidelity, with no kernel centering or positive-semidefinite correction applied.
B Model selection and search design
The scaling study uses common data handling and validation-F1 selection across methods while tuning each method on its own appropriate hyperparameter grid. Its broader grids adjust distance and interaction-phase scales for higher-dimensional settings and increasing interaction order.
- Common evaluation procedure: All methods use the same training–validation partition, validation-F1 selection criterion, complete-training-set refitting, and independent test evaluation.Each method selects its highest-validation-F1 configuration before one test-set evaluation.
- Common evaluation procedure: Each model searches a method-appropriate hyperparameter grid before selecting its final configuration.The selection procedure is shared, but the searched parameter grids are method-specific.
- Scaling search design: The higher-dimensional scaling study uses broad grids that adjust classical bandwidths and quantum interaction-phase ranges to appropriate scales.Classical bandwidths account for ambient dimension, while quantum phase ranges account for the decreasing characteristic magnitude of thin-slab block products as interaction order increases.
- Scaling search design: Regularization and remaining kernel parameters are tuned jointly over broad ranges.