Source-linked AI summary

Exploiting Nonlinear Recurrence and Fractal Scaling Properties for Voice Disorder Detection

Max A Little, Patrick E McSharry, Stephen J Roberts, Declan AE Costello, Irene M Moroz

arXiv:0707.0086v1nlin.CGnlin.CD

TL;DR

Voice disorders can profoundly affect patients, while speech biophysics includes nonlinear behavior that motivates broader analysis. The paper introduces combined recurrence and fractal-scaling measures, which distinguish normal subjects from subjects with all types of voice disorder, achieving 91.8 ± 2.0% overall correct classification performance.

  • Problem

    Voice disorders profoundly affect patients, and speech production includes inherent nonlinear biophysics requiring expanded analysis.

  • Method

    The paper introduces a combined nonlinear/stochastic signal analysis using recurrence and scaling methods.

  • Results

    91.8 ± 2.0% overall correct classification performance was achieved using the measures and noise components.

  • Takeaways & Limitations

    The two measures characterize aperiodicity and breath noise and distinguish normal subjects from subjects with all types of voice disorder.

  • Takeaways & Limitations

    The new measures rely upon sustained speech, and breathiness in speech is not usually affected in the same way.

Abstract

from arXiv · show

Voice disorders affect patients profoundly, and acoustic tools can potentially measure voice function objectively. Nonetheless, existing tools are limited to analysing voices displaying near periodicity, and do not account for inherent biophysical nonlinearity and non-Gaussian randomness. They do not directly measure complex nonlinear aperiodicity, and turbulent, aeroacoustic, non-Gaussian randomness. Often these tools have limited clinical usefulness. This paper introduces two new tools to speech analysis: recurrence and fractal scaling, which overcome the range limitations of existing tools by addressing directly these two symptoms of disorder, and a simple bootstrapped classifier distinguishes normal from disordered voices to 91.8% overall accuracy on a large database of subjects with a wide variety of voice disorders. They are widely applicable to the whole range of disordered voice phenomena by design. These new measures could therefore be used for a variety of practical clinical purposes.

Background

Existing acoustic measures are limited in their treatment of nonperiodic, nonlinear, and stochastic voice signals, motivating a unified framework that can characterize the full range of disordered voices. The paper introduces recurrence and fractal scaling measures within a deterministic–stochastic model and reports improved classification performance.

  • Voice disorders commonly involve increased vibrational aperiodicity and breath noise relative to normal voices.
  • Purely deterministic nonlinear methods cannot in principle characterize noise-dominated sounds that are better modeled as stochastic processes.This limitation prevents characterization of the full range of signals encountered clinically.
  • Existing perturbation and related measures do not directly quantify aperiodicity and breath noise or reliably reflect disorder severity.Their relationship with disorder extent or severity is not simple.
  • Most existing tools are properly applicable only when voices are near-periodic, leaving Type II and Type III sounds poorly characterized.Some algorithms produce no results for these sound types despite their clinical information.
  • The paper introduces a speech-production model that separates deterministic nonlinear and stochastic components, enabling methods to characterize both nonlinearity and randomness.The framework is intended to model dynamics across all types of disordered vowel speech.
  • The new measures achieve superior overall classification performance compared with classical perturbation measures and Michaelis’s derived irregularity and noise measures.

Methods

The paper models speech as arising from nonlinear vocal-fold dynamics, airflow, and stochastic fluctuations, then develops recurrence and fractal-scaling measures for irregularity and turbulent noise. These measures rank voices along aperiodicity and noise-related dimensions and support classification of healthy and disordered voices.

  • Biophysical model: Speech production involves nonlinear vocal-fold tissue dynamics coupled with airflow through the vocal tract.The governing description combines fluid dynamics with elastodynamics of a deformable solid.
  • Biophysical model: Aspiration noise arises from turbulent airflow and increases in some pathologies, contributing to perceived breathiness.Simulations indicate that incomplete vocal-fold closure can increase high-frequency noise.
  • Biophysical model: Classical linear acoustic and vocal-fold models exclude complicated turbulent airflow, although speech-relevant conditions can produce turbulence.The relevant Reynolds number is reported as very large, of order 10^5.
  • Stochastic model: Because point measurements of turbulence lose most fluid-dynamics information, a single measured variable can reasonably be modeled as a random process.The paper also notes that speech modeling should account for dynamical nonlinearity and randomness.
  • Signal measures: The analysis constructs recurrence and fractal-scaling measures to quantify irregular vibration and turbulent noise in voice signals.Higher Hnorm detects irregular vibration, while increased scaling exponent detects turbulent noise; normalized measures support ranking on aperiodicity and disorder-related scales.

Results

The results compare recurrence and DFA-derived measures with classical perturbation measures using hoarseness diagrams and classification tasks. The comparison uses 707 speech signals, while traditional perturbation values are available only for a smaller subset when their algorithms succeed.

  • Classification: The new-measure classification results are presented alongside direct comparisons using other combinations of classical perturbation measures.The comparison is organized through hoarseness diagrams and their associated classification performance results.
  • Experimental comparison: The comparison includes jitter, shimmer, and NHR as the three classical perturbation measures.Hoarseness diagrams also include combinations of these traditional measures alongside the new measures.
  • Experimental comparison: 707 speech signals were evaluated using normalized RPDE, DFA scaling exponents, and derived irregularity and noise components.The analysis compares these measures across the database of speech signals.
  • Experimental comparison: Traditional perturbation values were calculated for a smaller subject subset when the algorithms produced a result.The passage attributes details of these traditional algorithms to an external reference.
  • Classification: The best classification boundary was calculated using bootstrap resampling over 1000 trials.The classification results are summarized in Table 1 and applied to the hoarseness-diagram comparisons.

Discussion

The new nonlinear measures achieve stronger overall classification with fewer features than traditional approaches, while applying across speech-signal types by design. Their clinical use remains bounded by sustained-vowel and recording requirements, plus parameter-related reliability concerns.

  • Classification performance: 91.8 ± 2.0% overall correct classification is achieved by the RPDE/DFA pair.The jitter-and-shimmer combination produces the next-highest performance.
  • Classification performance: The new nonlinear measures are more accurate on average than traditional measures and derived irregularity and noise components under the same simple classifier.Matching the new measures’ performance would require more complex classifiers or many more classical features.
  • Feature complexity: Only two new measures are required for good separation performance, helping mitigate the curse of dimensionality.Increasing feature count expands feature-space volume exponentially, while the available training examples occupy an increasingly small volume.
  • Feature complexity: Traditional measures show high positive correlation and occupy an effectively one-dimensional object, whereas the new measures are spread evenly across the feature space.The irregularity and noise components occupy more of the feature-space area than traditional measures.
  • Parameterisation: The new measures require five arbitrary parameters and can calibrate out dependence on recurrence-radius and state-space reconstruction parameters using analytical results.Parameter selection still matters because specialised settings may separate one dataset well without generalising to new data.
  • Scope and clinical limitations: The new measures are capable by design of measuring all types of speech signals, unlike traditional measures whose coherent interpretation breaks down for Type II/III and random-noise signals.In clinical practice, the new measures rely on sustained vowel phonation, may require discarding the beginning of phonation, and require a constant microphone distance.

Conclusions

The paper combines recurrence and scaling analysis to characterize nonlinear aperiodicity and non-Gaussian breath noise across normal and disordered voices. Under sustained-vowel, quiet-recording conditions, the two measures distinguish normal from disordered subjects and outperform existing approaches while using fewer algorithmic parameters.

  • The study targets aperiodicity and breath noise as the two main biophysical factors of disordered voices.
  • The authors introduce a combined nonlinear/stochastic signal model and explore recurrence period density entropy with detrended fluctuation analysis.
  • The model is intended to produce the wide variation in behavior observed across normal and disordered voice examples.
  • Under sustained-vowel recordings made in quiet acoustic conditions, the measures directly characterize the targeted disorder-related factors.
  • The two measures alone distinguish normal subjects from subjects with all types of voice disorder, with better classification performance than existing measures.
  • Existing approaches are considerably more complex when their arbitrary algorithmic parameters and required measure combinations are considered.
  • The authors conclude that speech-production nonlinearity and non-Gaussianity can support signal-analysis methods and screening systems better able to characterize voice-disease variation.

Periodic Recurrence Probability Density

This section derives first-recurrence probabilities for a purely deterministic, finite-period trajectory in reconstructed state space. Points on the periodic orbit recur with certainty after the period, while other points never recur; the whole-space probability sums contributions from the orbit points.

  • The derivation considers the purely deterministic case, where the speech-production model has no forcing term ε(t).
  • A finite-period trajectory repeats after k steps, with distinct states within each period.
  • For a point on the periodic orbit, the first return occurs with certainty after k time steps.
  • A point not belonging to the periodic orbit is never reached by the trajectory and therefore has no first return.
  • The first-recurrence probability for the whole reconstructed space is obtained by summing appropriately weighted probabilities for the individual points.

Uniform i.i.d. Stochastic Recurrence Probability Density

For a uniform i.i.d. stochastic signal, trajectories occupy a partitioned reconstructed state space, and the first-recurrence probability density is uniform over the measured recurrence-time range.

  • A uniform i.i.d. forcing term produces a stochastic, i.i.d. measured trajectory.
  • The normalised signal reconstructs trajectories in the state space [−1, 1]^m, partitioned into N^m equal-sized cubes.Each cube has side length Δs = 2/N and probability P_R = Δs^m/2^m.
  • Recurrence to the whole space is obtained by summing the appropriately weighted first-recurrence probabilities for the individual cubes.
  • For small cube sizes and close-return radius, the close-returns algorithm determines the first-recurrence probability over a finite range 1 ≤ T ≤ T_max.
  • The recurrence probability density is uniform for a uniform i.i.d. stochastic signal.

Authors’ Contributions

The authors divided responsibilities across study design, mathematical methods, software and data analysis, discriminant analysis, data preparation, and manuscript approval.

  • MAL led the conceptual design, developed the mathematical methods, wrote the software and data-analysis tools, and prepared and analysed the data.
  • PEM contributed to the conceptual design and mathematical methods, while SJR participated in discriminant analysis.
  • DAEC contributed to data preparation, and IMM contributed to developing the mathematical methods.
  • All authors read and approved the manuscript.

Table 1 - Summary of disordered voice classification results

The RPDE/DFA combination produced the strongest reported classification results among the listed measure combinations, outperforming the traditional perturbation combinations on overall accuracy.

  • 91.8±2.0% overall accuracy was achieved by RPDE/DFA across 707 subjects.Its true-positive rate was 95.4±3.2% and true-negative rate was 91.5±2.3%.
  • 81.4±3.9% overall accuracy was achieved by Jitt/Shim across 685 subjects.
  • 80.7±4.0% overall accuracy was achieved by Shim/NHR across 684 subjects.
  • 79.3±5.5% overall accuracy was achieved by Irreg/Noise across 707 subjects.
  • 76.4±4.8% overall accuracy was achieved by Jitt/NHR across 684 subjects.

Additional Files

The supplied additional-file materials include implementations of the close-returns and DFA algorithms, along with figures illustrating signal embedding, recurrence analysis, RPDE, scaling analysis, and classification.

  • The close-returns algorithm was implemented in C with a Matlab MEX interface and compiled as a Windows DLL.
  • The DFA algorithm was implemented in C with a Matlab MEX interface, compiled as a Windows DLL, and wrapped in a Matlab function.
  • Figure 2 illustrates time-delay embedding of normal and disordered speech signals using m = 3 and τ = 7 samples.
  • Figure 3 demonstrates recurrence analysis on perfectly periodic and uniform i.i.d. random signals.
  • Figure 4 compares RPDE for normal and disordered voices, while Figure 5 presents scaling-analysis examples.
  • Figure 6 presents hoarseness diagrams comparing new and classical measures for normal/disordered classification.
Loading 0707.0086v1…