Source-linked AI summary
TRACE-C: Rank-Calibrated Relational Anomaly Detection for Multi-Stream Operational Telemetry
Matthew Faucher
TL;DR
Operational telemetry can be anomalous through joint configurations even when individual streams look familiar. TRACE-C is an auditable strictly-prior rank-calibrated detector that combines local, dependence, and temporal channels, but its mixed evaluation results and disclosed calibration limits narrow the supported conclusion to an evidence-bound operating envelope.
Problem
Operational telemetry can become jointly anomalous while each individual stream remains within its familiar range, making multivariate detection necessary.
Method
TRACE-C applies strictly-prior online scoring to aligned multi-stream telemetry using rolling robust-z residuals, three window channels, rank aggregation, and explicit selection rules.
Results
TRACE-C ranks Storm Atiyah first in 2019, but ablation shows a local-led result; fusion ranks the short frequency event below the temporal channel, while 2020 selects no alerts.
Takeaways & Limitations
The evidence supports TRACE-C as an auditable algorithm and evidence package whose assumptions, fallbacks, trade-offs, and negative results are reported alongside its successes.
Takeaways & Limitations
Rank counts are descriptive because exchangeability, coverage, and FDR control are not established for the reused dependent online references and record fallback.
Abstract
from arXiv · showhide
Operational telemetry can be jointly anomalous while every individual stream stays inside its familiar range. TRACE-C is an auditable strictly-prior rank-calibrated detector for aligned multi-stream telemetry: same-regime rolling median/MAD residuals feed three window channels -- a maximum normalized local sum, a Gaussian copula-form dependence contrast on robust-z residuals, and a worst standardized AR(1) innovation -- whose channel ranks are Fisher-aggregated and ranked against earlier aggregates. We evaluate six Great Britain grid streams with a January-April 2019 fit, July-December 2019 development evidence, and a 2020 hold-out frozen before inspection. TRACE-C ranks Storm Atiyah first among 2019 test windows, but a disclosed channel ablation attributes that rank to the local channel, not the copula-form channel: copula-only ranks Atiyah 59th. The short 9 August frequency event is ranked far lower by the fused detector (143) than by the temporal channel alone (40), and reconstruction baselines rank it first. In 2020 no window is selected, which is consistent with record-rule saturation rather than an uneventful year; the highest-ranked frozen window was later interpreted as Storm Ellen. Three interpretive limits carry throughout. The resulting p-values are selection quantities, not event probabilities. The copula-form channel is not a literal copula density: the method applies no probability-integral or normal-score transform. Empirical rank counts are diagnostics, not coverage or false-discovery proofs. Every table and figure in this paper is generated from committed machine-readable reports.
1 Introduction
TRACE-C addresses multivariate operational anomalies that may be invisible in individual streams by combining strictly-prior rank-calibrated channels with explicit selection and evidence reporting. The paper reports mixed results and limits its claims to an auditable operating envelope rather than validated copula discovery or error control.
- Operational windows can be jointly unusual even when each stream remains within its familiar range, while regime drift can inflate conventional anomaly scores without a discrete event.
- TRACE-C uses an auditable three-channel detector whose references are strictly prior to each scored window.
- The implemented dependence channel has Gaussian copula-form algebra but is not literal copula calibration, and Fisher aggregation is distinguished from a chi-square test.
- The selection path attempts ordinary BH, discloses a record-rule fallback, and includes a fixed-block alert budget whose assumptions do not support a nominal FDR claim.
- Storm Atiyah ranks first in 2019 development, but ablation identifies a local-led result and Fisher fusion can bury the short frequency event relative to the temporal channel alone.
- The evaluation uses a frozen 2020 hold-out, checksummed retained inputs, and machine-readable reports to generate every table and figure.
2 Related work
The related-work context distinguishes rank-based guarantees and dependence modeling from TRACE-C’s reused online reference scheme. It also positions the comparison against established anomaly-score families while emphasizing that the implemented copula-form algebra and selection procedures do not inherit standard calibration guarantees.
- Dependent telemetry can invalidate exchangeability-based rank interpretations, motivating block permutations or related methods under explicit dependence conditions.
- Reference/test results for conformal outlier testing do not validate TRACE-C’s growing, reused online reference, so its rank counts remain empirical diagnostics.
- TRACE-C applies a Gaussian copula quadratic contrast to robust-z residuals without the marginal transform, making “copula-form” an algebraic description rather than a fitted copula density.
- Fisher aggregation is used only as a score and is recalibrated by online rank because chi-square calibration conditions are not asserted for dependent channel ranks.
- BH’s standard FDR guarantees require dependence conditions not established here, so TRACE-C uses BH as a nominal selection attempt rather than demonstrated FDR control.
- The record fallback has an exchangeable-continuous expected-count benchmark, but serial dependence, ties, and drift prevent that count from becoming a present-application guarantee.
- The post-hoc baseline comparison covers autoencoder, PCA, Isolation Forest, and Spectral Residual anomaly-score families.
3 Method
TRACE-C performs strictly-prior online scoring on aligned telemetry, transforming same-regime robust residuals into three complementary window channels. Channel ranks are Fisher-aggregated and then ranked against strictly earlier aggregates, with explicit selection and attention-budget rules.
- Windowing and residuals: TRACE-C partitions aligned telemetry into non-overlapping four-row windows and permits only strictly prior observations or fitted parameters when scoring each window.The four-row windows correspond to two hours on ordinary 48-period days.
- Windowing and residuals: Each stream and settlement-period/weekday regime uses the 40 most recent prior values to compute rolling median/MAD robust residuals.A small positive floor replaces zero MAD in code.
- Three channels: The local channel takes the largest absolute normalized window sum, preserving sustained magnitude in at least one stream without concealing the responsible sensor.The maximum favors sustained magnitude while avoiding a cross-stream sum that could hide which sensor contributes.
- Three channels: The relational channel uses a Gaussian copula-form contrast based on robust residuals, but applies neither a probability-integral transform nor a normal-score transform.It is therefore not a literal copula density; the name describes its algebraic form rather than calibration.
- Three channels: The temporal channel uses per-stream AR(1) parameters and innovation scales fitted on the designated fit segment, taking the worst standardized innovation in each window.A maximum preserves a sharp transition that a mean could dilute across the two-hour window.
- Rank calibration: Each channel is ranked against its trailing 240 strictly prior scores, Fisher aggregation combines the channel ranks, and the aggregate receives an outer strictly-prior rank after 40 prior aggregates.The Fisher expression is only an aggregation score, without a claimed chi-square null distribution.
- Selection and budget: TRACE-C applies nominal BH at q = 0.05, falls back to an upper-record rule when BH selects nothing, and retains at most two windows per fixed 48-row block.The block limit is an attention budget rather than a calendar-day or daylight-saving-safe construction.
4 Data and evaluation protocol
The evaluation uses six Great Britain electricity-system telemetry streams with pinned inputs, a staged 2019 development chronology, and a frozen 2020 hold-out. Event annotations and baseline comparisons follow distinct ranking protocols, while committed reports support reproducible arithmetic without guaranteeing bitwise-identical artifacts or statistical coverage.
- Data: Six streams combine five half-hour demand-side or generation measurements with a frequency-deviation summary derived from NESO telemetry.The demand-side columns are national demand, transmission-system demand, embedded wind, embedded solar, and pumped-storage pumping; frequency is summarized by maximum absolute deviation from 50 Hz per settlement period.
- Chronology: January–April 2019 fits the dependence matrix and AR(1) parameters, May–June supplies prior windows, and July–December provides 2,208 development windows.The fit segment is assumed representative enough for parameter fitting but is not established to be anomaly-free; W = 4 and K = 40 were informed by 2019 evidence.
- Chronology: All 4,392 scored 2020 windows form a hold-out whose method, sensor set, and configuration were fixed before inspection.The hold-out is frozen by development chronology, not treated as a clean null year, because 2020 includes major weather and demand-regime changes.
- Event evaluation: Event annotations are joined only after ranking, with intervals spanning 5 windows for the 9 August disturbance, 12 for Atiyah, and 24 each for Ciara and Dennis.The 2020 lockdown interval covers 168 windows; Ellen and Alex were named retrospectively and were not detector inputs or pre-specified evaluation intervals.
- Baselines: Four baselines share the source rows, six-stream sensor set, windows, cutoff, dates, score direction, and event intervals, but their ranking and selection protocol differs from TRACE-C.Baselines use global fit scaling, raw-score strictly-prior ranks after warm-up, and an unbudgeted record rule; TRACE-C uses robust residuals, channel and outer ranks, Fisher aggregation, BH-first selection, and a two-per-block budget.
- Reproducibility: Pinned inputs, hashes, runtime metadata, and JSON-derived tables support computational reproduction of the reported arithmetic, while empirical rank-p expected counts are not time-series coverage guarantees.The reports do not claim that external-model artifacts will be bitwise-identical across platforms.
5 Results
TRACE-C’s results are mixed across events, channels, years, and detector families. Storm Atiyah ranks first in 2019, the short frequency event is better captured by temporal or reconstruction methods, and no 2020 window is selected.
- Empirical rank diagnostics: 120/110.4 at .05 and 23/22.1 at .01 are the 2019 empirical rank counts versus arithmetic benchmarks, while pre-COVID 2020 reports 50/45.0 and 3/9.0.The full 2020 report gives 205/219.6 at .05 and 39/43.9 at .01; these are descriptive diagnostics.
- Selection outcomes: Two 2019 development windows were selected versus an exchangeable-continuous benchmark of 1.19, while zero 2020 windows were selected versus 0.87.BH selected no windows in either segment, so the disclosed record fallback ran; the zero-alert 2020 result is consistent with record saturation.
- 2019 event results: The 9 August power-cut interval ranks 143 with p = 0.060420 across 5 opportunities, despite the frequency-deviation stream.The event was operationally important, but importance and separability at the chosen aggregation differ.
- 2019 channel ablation: Atiyah ranks 59 under the copula-form channel alone, while the short frequency event ranks 40 temporally but 143 when fused.Dropping the copula channel leaves Atiyah at rank 2 without a record-rule alert; local-only gives rank 3, and dropping local removes alerts.
- Frozen 2020 hold-out: In 2020, Ciara’s best window ranks 44, Dennis’s 137, and the best lockdown window 45; the highest-ranked window was later interpreted as Storm Ellen.The lockdown result is not onset detection, and the post-ranking storm interpretations are not detector discoveries.
- Baseline comparison: The convolutional autoencoder and PCA rank the 2019 frequency event first, while Isolation Forest and Spectral Residual rank it second.For Atiyah, TRACE-C ranks first while those baselines rank it 34, 12, 3, and 26, respectively.
- Baseline comparison: No detector dominates: outcomes vary with score construction, interval length, sensor set, and temporal resolution.The comparison is exploratory rather than a universal ranking of algorithms.
6 Limitations and honesty ledger
The paper bounds its claims around dependence, post-hoc choices, aggregation, governance, and deployment. Its rank arithmetic and record fallback are transparent, but the evidence does not establish coverage, FDR, or broad operational applicability.
- Calibration under dependence: TRACE-C’s reused, adaptive references operate on telemetry not shown to be exchangeable, so empirical rank counts are descriptive and coverage or FDR is not established.The growing outer reference reuses history across tests, and ordinary BH remains nominal.
- Record-rule limitation: The record fallback can saturate over long horizons: an early extreme may suppress later records, as seen with zero selected alerts in 2020.Its expected count is an exchangeable-continuous benchmark rather than a time-series guarantee.
- Relational terminology: The Gaussian copula-form channel is not a literal copula density because robust-z inputs lack the probability-integral transform, and Fisher’s chi-square reference is not used.The channel is treated as a fitted dependence contrast, while the aggregate is recalibrated by online rank.
- Detectability: Two-hour non-overlapping windows and a half-hour maximum frequency deviation can erase short-transient timing or let sustained background extremes dominate.Aggregation and sensor selection therefore define detectability.
- Baseline comparison: The baseline comparison is post-hoc and protocol-nonidentical, with only TRACE-C frozen before 2020 inspection; apparent wins are exploratory.Baselines use global scaling, raw-score ranks, and unbudgeted record selection rather than TRACE-C’s regime conditioning and fixed-block budget.
- Deployment scope: Only two windows may be selected per fixed 48-row block, which is not a calendar service across 46- or 50-period days or daylight-saving transitions.Deployment would require a calendar-aware budget and policy for late or revised telemetry.
- Governance scope: TRACE-C is scoped to process and sensor anomalies for human review, not person-risk scoring or decisions about individuals.The study’s evidence cannot support individual-level decisions.
7 Conclusion
TRACE-C is an auditable strictly-prior rank-calibrated detector whose evidence exposes fusion trade-offs, record saturation, and sensor or aggregation limits rather than establishing nominal FDR or copula discovery.
- TRACE-C combines magnitude-preserving robust residuals, local, dependence, and temporal channels, rank aggregation, and an explicit selection path.
- Storm Atiyah ranks first in 2019 development, but copula-only ranks it 59th and removing copula leaves rank 2 without a record alert.
- The short 2019 frequency event is ranked better by simple reconstruction and the temporal channel alone than by the fused detector.
- The frozen 2020 run selects no alerts, while its ranked list contains plausible weather-associated and unlabelled extremes.
- Future evaluation should use a pre-registered dependence-aware selector, untouched data, sub-period frequency information, alternative aggregations, and calendar-aware budgeting.
Code and data availability
The implementation and reproduction materials accompany the manuscript, with checksummed telemetry inputs and provenance for generated reports.
- The implementation, committed machine-readable reports, table generator, and reproduction instructions accompany the manuscript.
- Retained telemetry inputs are covered by checksums and report-generator provenance.
- The data are provided by the National Energy System Operator under the NESO Open Data Licence.