Source-linked AI summary

Highly comparative time-series analysis: The empirical structure of time series and their methods

Ben D. Fulcher, Max A. Little, Nick S. Jones

arXiv:1304.1209v1physics.data-ancs.CVphysics.bio-phq-bio.QMstat.ML

TL;DR

Scientific time-series data and analysis methods have not been extensively organized across disciplines, making their relationships and suitable applications difficult to assess. The paper builds empirical fingerprints from large annotated libraries, representing time series by measured properties and methods by behaviour across data. These representations organize interdisciplinary resources, retrieve comparable alternatives, and support automated analysis-method selection across several datasets.

  • Problem

    The number and interdisciplinary diversity of time-series methods and datasets make their relationships and appropriate use difficult to determine.

  • Method

    The paper represents time series by properties measured with diverse operations and represents operations by outputs across 875 interdisciplinary time series.

  • Results

    The representations organize time-series datasets and methods, retrieve similar real-world and model-generated data or alternative methods, and support automated selection for classification and regression.

  • Takeaways & Limitations

    Highly comparative analysis provides an interdisciplinary context that can guide more focused time-series analysis and connect methods and data across scientific fields.

Abstract

from arXiv · show

The process of collecting and organizing sets of observations represents a common theme throughout the history of science. However, despite the ubiquity of scientists measuring, recording, and analyzing the dynamics of different processes, an extensive organization of scientific time-series data and analysis methods has never been performed. Addressing this, annotated collections of over 35 000 real-world and model-generated time series and over 9000 time-series analysis algorithms are analyzed in this work. We introduce reduced representations of both time series, in terms of their properties measured by diverse scientific methods, and of time-series analysis methods, in terms of their behaviour on empirical time series, and use them to organize these interdisciplinary resources. This new approach to comparing across diverse scientific data and methods allows us to organize time-series datasets automatically according to their properties, retrieve alternatives to particular analysis methods developed in other scientific disciplines, and automate the selection of useful methods for time-series classification and regression tasks. The broad scientific utility of these tools is demonstrated on datasets of electroencephalograms, self-affine time series, heart beat intervals, speech signals, and others, in each case contributing novel analysis techniques to the existing literature. Highly comparative techniques that compare across an interdisciplinary literature can thus be used to guide more focused research in time-series analysis for applications across the scientific disciplines.

1 Introduction

Time-series analysis spans many scientific disciplines, but the diversity of methods and data makes cross-disciplinary comparison and method selection difficult. The paper addresses this gap by jointly organizing annotated libraries of time series and analysis methods through their empirical properties and behaviour.

  • Time series are fundamental data objects across disciplines, including finance, astrophysics, meteorology, and medicine.
  • Methods such as fluctuation analysis, GARCH models, and Sample Entropy have developed in different disciplinary contexts.
  • The paper assembles annotated libraries and uses methods’ behaviour across data, and time series’ measured properties, to create a unified platform for comparison.
  • The number and interdisciplinary diversity of methods make it difficult to determine how methods relate or which are appropriate for finite, noisy data.
  • It is similarly difficult to compare time series across scientific disciplines or with dynamics generated by theoretical models.

2 Framework

The framework combines large annotated libraries of time series and analysis algorithms in a data matrix. Rows fingerprint time series by measured properties, while columns fingerprint operations by their outputs across time series.

  • 38 190 univariate time series and 9 613 time-series analysis algorithms form the paper’s annotated libraries.
  • Each analysis method is implemented as an operation that maps an input time series to a single real-number summary.
  • The operation library quantifies diverse properties, including distributional statistics, linear correlations, stationarity, and other dynamics.
  • Parameter variants can make the library exceed the number of conceptually distinct methods, estimated at approximately 1 000 unique operations.
  • The matrix element D_ij equals the output of operation F_j applied to time series x_i, D_ij = F_j(x_i).
  • The matrix exposes relationships among measurements and systems, including redundancy when operations produce similar output patterns across time series.

3 Empirical structure

Empirical fingerprints organize both analysis methods and time series across a large interdisciplinary collection. The resulting structure supports redundancy reduction, method retrieval, dataset clustering, and automated links between real-world signals and model systems.

  • Empirical structure of time-series analysis methods: 8 651 operations were organized by clustering their behaviour across 875 interdisciplinary real-world and model-generated time series.
  • Empirical structure of time-series analysis methods: 200 operations approximate the full 8 651-operation library with residual variance 0.05, exploiting redundancy among methods.
  • Empirical structure of time-series analysis methods: Empirical fingerprints retrieve alternatives to a target method, including methods from unfamiliar fields with similar behaviour across data.
  • Empirical structure of time series: 24 577 time series were represented by 200 measured properties and clustered into 2 000 groups after filtering the larger library.
  • Empirical structure of time series: The time-series representation retrieves real-world and model-generated neighbours, linking Oxford Instruments share prices to stochastic differential equation models.
  • Empirical structure of time series: Targeting a stochastic sine map retrieved meteorological series with qualitatively similar noisy-switching dynamics.
  • Together, 200 operations and 875 time series were sufficient to organize methods and data meaningfully across the analyzed collections.

4 Applications

The highly comparative framework organizes diverse time series and analysis operations by empirical behaviour, then uses these representations to select useful methods for classification and regression. Case studies show that simple, interpretable operations can organize EEG and heart-rate methods, identify novel or redundant measures, and provide alternatives for estimating known signal properties.

  • Framework: The framework organizes time-series datasets and analysis methods using measured properties and empirical outputs, respectively.It supports organizing datasets, retrieving useful operations, and selecting methods for classification and regression tasks.
  • 4.1 EEG recordings: 172 operations individually classified healthy EEGs and seizures with cross-validation rates above 95%, including eight exceeding 98.75%.These operations came from diverse analysis literatures and provided interpretable differences in entropy, correlation dimension, scaling exponents, and distributional properties.
  • 4.2 Heart rate variability: 60 successful heart-rate operations clustered into three main behavioural groups: location, entropy/complexity, and linear-correlation measures.The matrix also distinguished a wavelet operation and an outlier-adjusted autocorrelation measure with relatively unique behaviour.
  • 4.2 Heart rate variability: Comparative analysis can distinguish redundant methods from genuinely novel operations in an analysis literature.The variance ratio hypothesis test reproduced existing linear-model behaviour, whereas an outlier-adjusted autocorrelation operation was both unique and useful for the dataset.
  • 4.3 Self-affine time series: Five operations showed strong linear relationships with the scaling exponent α, including fluctuation-analysis, autoregressive, forecasting, and Gaussian-process methods.The alternatives captured α through combinations of local and global structure and could require less computation or support iterative updates.
  • 4.4 Parkinsonian speech: The approach selects interpretable operations and complementary feature pairs for classification across scientific recordings.Applications included Parkinsonian speech, seismic classification, and seven-class emotional speech, while simple statistical learning methods produced meaningful results.

5 Conclusions

The study shows that empirical comparisons can reveal meaningful structure in large interdisciplinary collections of time series and analysis methods. This framework gives analysts tools to contextualize methods, compare datasets, select informative operations, and guide more focused research.

  • 5 Conclusions: Empirical behaviour provides a basis for organizing extensive collections of time-series data and analysis methods.Methods are represented by outputs on data, while time series are represented by properties measured through operations.
  • 5 Conclusions: The framework lets analysts compare familiar methods with alternatives from other disciplines and check whether new methods reproduce existing behaviour.These comparisons can reveal redundancy and help assess whether proposed methods constitute advances.
  • 5 Conclusions: Structuring datasets and connecting them to real-world and model systems provides interdisciplinary context for more focused analysis.The resulting organization relates datasets to relevant types of systems and their measured properties.
  • 5 Conclusions: Selecting operations according to their behaviour yields interpretable information about which properties distinguish labeled classes.The selected operations connect informative data properties to automated method selection.
  • 5 Conclusions: Highly comparative analysis complements domain-specific methods as scientific time-series data and analysis techniques continue to expand.The paper presents this comparison across interdisciplinary literatures as analogous to high-throughput analysis guiding focused research in biology.
Loading 1304.1209v1…