Source-linked AI summary

From local kernels to global form: modeling the emergence of musical content

Francesco Vitucci, Michele Lorusso, Francesco Scagliola

arXiv:2608.24660v1cs.CL

TL;DR

The paper addresses the limitation of a single globally constant transition kernel for musical sequences whose syntactic law changes over time. It derives local directed transition kernels from overlapping windows of one symbolic sequence and evaluates them analytically and generatively. In Debussy’s Syrinx, both reference boundaries align across pitch and duration at L = 6, but geometry and broad plateaus limit claims of unique automatic segmentation.

  • Problem

    A time-homogeneous first-order model treats occurrences identically across formal regions, compressing long-range structural information into one globally constant kernel.

  • Method

    Overlapping sliding windows derive a temporal trajectory of directed conditional-transition kernels from one symbolic sequence, evaluated in pitch and duration through distance tests and re-synthesis.

  • Results

    At L = 6, both score-based references coincide with maximal change in pitch and duration, while the rhythmic trajectory is more selective and re-synthesis shows substantial event-level departure at L = 6 and L = 20.

  • Takeaways & Limitations

    Local kernels provide boundary-aligned, multiscale evidence consistent with formal change, but broad plateaus prevent treating either distance curve as a unique automatic segmenter.

  • Takeaways & Limitations

    The strongly overlapping distance compares one entering and one leaving transition, making it sensitive to local turnover but limiting interpretation as a true boundary detector.

Abstract

from arXiv · show

Markov models are established tools for symbolic music, including non-homogeneous formulations. The narrower contribution examined here is an observation-driven estimation mechanism: overlapping sliding windows derive a trajectory of local transition kernels from one symbolic sequence rather than from an exogenous formal partition. We test this mechanism on 273 logical note events from Debussy's Syrinx (1913), using the often-proposed A-B-A' reading as a reference rather than ground truth. We apply the same validation to absolute-pitch and notated-duration kernels. At $L=6$, both reference boundaries attain the Jensen--Shannon maximum in both dimensions; the duration plateau is substantially narrower (64 of 267 comparisons) than the pitch plateau (210 of 267). Because the theoretical maximum for consecutive sliding-window comparisons is set by window geometry and equals $1/\sqrt{L-1}$ for maximal turnover of the entering/leaving transition, the pitch value at $L=6$ and its broad plateau are not, by themselves, strong evidence. Their cross-dimensional alignment is consistent with boundary sensitivity, while the broad plateaus preclude treating either curve alone as a unique automatic segmenter. Five-hundred-draw re-synthesis experiments quantify departure from the source in both dimensions and expose an exact-copy degeneracy at $L=2$.

1. INTRODUCTION

Markov models provide transparent tools for capturing local musical regularities, while this paper studies observation-driven local transition kernels derived from one symbolic sequence. It asks whether directed windowed kernels can support analytical or generative evidence without relying on an exogenous formal partition.

  • Markov processes have long supported symbolic-music applications because transition matrices capture local regularities and remain inspectable.They have been used in algorithmic composition, style imitation, improvisation, and hybrid symbolic/audio modelling.
  • Time-varying transition probabilities relax time-homogeneity while preserving first-order memorylessness.The next state still depends only on the current state, but the transition law may change with position.
  • The contribution combines directed conditional-transition structures, their temporal trajectory, and analytical and generative uses from one symbolic sequence.The paper distinguishes this combination from non-homogeneity, sliding-window estimation, and local statistics considered separately.
  • Sliding-window estimation motivates testing whether directed conditional transitions provide useful analytical or generative evidence in a single-piece case study.This extends windowed local analysis beyond marginal pitch-class histograms.

2. BACKGROUND: MARKOV MODELS

The background formalizes first-order Markov chains, transition matrices, and empirical estimation for symbolic event sequences. The implementation represents sequence boundaries explicitly while excluding the terminal row from generative sampling.

  • A first-order Markov chain makes the next state conditionally independent of all earlier states given the current state.The state space is finite, and musical states may represent pitches, pitch classes, chords, durations, or composite tokens.
  • A time-homogeneous chain uses one transition matrix P whose rows define conditional next-state probabilities.Each row is normalized as a probability distribution over destination states.
  • Empirical estimation counts observed transitions Cij and normalizes each source-state row.The estimator is applied to an observed symbolic sequence.
  • Boundary tokens ⟨START⟩ and ⟨END⟩ explicitly represent first and last events as transitions.The terminal row has no outgoing mass and is excluded from probability sampling during generation.

3. WHY ESTIMATE LOCAL KERNELS?

A single globally constant transition kernel compresses long-range musical structure and treats identical transitions as equivalent across formal regions. The paper therefore motivates kernels that preserve memorylessness while representing temporal location and changing syntactic laws.

  • A globally constant first-order kernel compresses all long-range structural information into one matrix.The paper separates the Markov property from the additional assumption of time-homogeneity.
  • Musical events gain meaning through temporal position, recurrence, expectation, tension, and resolution rather than occurring as isolated symbols.The motivation is to retain temporal structure in the transition law.
  • Homogeneous models cannot distinguish transitions occurring in structurally distinct regions when those regions are perceptually or syntactically dissimilar.This limitation concerns time-invariance, not first-order conditional independence itself.
  • Variable-order models capture depth of context, whereas the proposed approach captures location in time.Higher-order models can also produce rapid state-space growth and data sparsity in finite or windowed corpora.

4. TIME-VARYING MARKOV MODEL VIA SLIDING WINDOWS

The method estimates overlapping local transition matrices from short symbolic windows, producing a trajectory that represents changing musical grammar and supports time-inhomogeneous generation.

  • Sliding Window Construction: Overlapping windows of length L contain L−1 transitions and yield a sequence of local count matrices for successive temporal segments.With step size 1, window m uses transitions indexed by I_m = {m, ..., m + L −2}.
  • Sliding Window Construction: Each local matrix is interpreted as a locally stationary kernel under the assumption that transition probabilities are approximately constant within its window.The window length controls a bias–variance trade-off: larger windows reduce variance but blur change, while smaller windows track change with sparser rows.
  • Implementation Details: For analytical comparisons, boundary tokens are masked, while generation uses row-wise fallback because an observed local matrix may not itself be a complete stochastic kernel.The baseline uses unsmoothed counts and backs off to the global kernel only when the selected local row is empty.
  • Time-Varying Transition Kernels: The resulting trajectory P = {bP (m)} records local stylistic behaviour, temporal evolution of musical grammar, and a piecewise-constant approximation to a time-inhomogeneous process.Successive windows overlap by L−1 symbols, so the estimated kernel changes smoothly as the window advances.
  • Time-Varying Transition Kernels: The model retains first-order memorylessness while relaxing time-homogeneity, allowing transition probabilities to vary with position in the sequence.A schedule g(r) maps each transition time to one local estimate, producing a time-inhomogeneous Markov chain.
  • Generative Use: Generative use selects local pitch and duration kernels on a shared schedule, keeping the two streams separately inspectable rather than forming a joint pitch–rhythm state space.Centre alignment is used for the schedule, with global-kernel fallback when a selected local row has zero mass.

5. CASE STUDY: DEBUSSY’S SYRINX

The Syrinx case study derives local pitch and duration transition kernels from overlapping windows and compares their successive changes with A–B–A’ reference coordinates. At L = 6, both references reach the geometric JS maximum in both dimensions, while re-synthesis evaluates departure from the source.

  • Data and setup: Measures 9 and 26 serve as falsifiable A–B–A’ comparison references rather than ground-truth sectional labels.Published analyses converge on the tripartite macroform but disagree over internal divisions.
  • Data and setup: 273 aligned note events from Syrinx provide absolute-pitch and exact-notated-duration streams after excluding grace notes and merging tied continuations.The score contains 308 written note entries; 14 zero-duration grace notes are excluded and 21 tied continuations merged.
  • Distance test: At L = 6, both references attain the maximum JS distance of 0.447 in both pitch and duration streams.The comparison uses successive transition-distribution distances between overlapping windows.
  • Distance test: The L = 6 maximum is geometrically expected, while duration is more selective: 64 of 267 comparisons form its plateau versus 210 of 267 for pitch.The theoretical ceiling is 1/sqrt(L-1) under maximal turnover of the entering and leaving transition.
  • Generative re-synthesis: Five-hundred-draw re-synthesis evaluates aligned mismatch and transition JS for pitch and duration, with L = 2 producing exact copies of the source.Hard backoff is retained as the explicit baseline; shrinkage removes the degeneracy but is not neutral.

6. DISCUSSION

The discussion treats the distance test as bounded case-study evidence: cross-dimensional alignment at L = 6 is consistent with boundary sensitivity, but overlapping-window geometry and broad plateaus limit automatic segmentation claims.

  • Interpretation: At L = 6, pitch and duration references coincide with maximal change, with the rhythmic trajectory showing greater selectivity.The result is boundary-aligned evidence from two complementary event dimensions.
  • Interpretation: Broad plateaus mean descriptive validation does not constitute a turnkey segmentation algorithm.Published analyses’ differing internal partitions support a multiscale reading rather than one mandatory partition.
  • Limitations: Strongly overlapping consecutive windows measure local turnover, so their geometry limits interpretation as a true boundary detector.Future work should compare boundary-centred preceding and following windows, preferably non-overlapping.
  • Generative implications: Generative evaluation should distinguish memorisation, event-level departure, within-section syntax, pitch–duration coherence, and perceptual judgement.Formal controllability cannot be inferred from a single piano roll or global distributional similarity.

7. RELATED WORK

Prior work establishes homogeneous and non-homogeneous Markov models, sliding-window statistics, and deep-learning approaches for symbolic music. This paper instead evaluates window-derived directed transition structures and their temporal trajectory within a single symbolic sequence.

  • Markov models: Most existing Markov approaches use homogeneous models, while prior non-homogeneous models condition transitions on hand-specified phrases or recurring metrical position.These approaches differ from conditioning on local position inferred directly from one piece.
  • Windowed local statistics: Sliding-window estimation is established in nonstationary analysis and MIR, but earlier musical work used marginal pitch-class histograms rather than directed conditional transitions.The distinction is between local distributions over events and local transition kernels.
  • Paper’s position: The contribution is applying and evaluating window-derived directed transition structures and their temporal trajectory for analysis and generation in a single symbolic sequence.The paper explicitly distinguishes this combination from sliding windows, local statistics, and non-homogeneous chains individually.
  • Relation to context models: Variable-order models address context depth, whereas this approach addresses temporal location; the two relaxations are in principle combinable.This frames the method as complementary to variable-order modeling rather than as a replacement.
  • Deep-learning comparison: Deep-learning systems achieve strong symbolic-music-generation results but require large datasets and do not provide interpretable, time-resolved transition matrices.The paper characterizes the paradigms as complementary rather than competing.

8. CONCLUSION

The study presents overlapping windows as an interpretable way to track changing transition kernels and tests their sensitivity to formal boundaries in Syrinx. The findings support boundary sensitivity in this first case without establishing unique segmentation or corpus-level generality.

  • Conclusion: At L = 6, both reference boundaries attain maximal successive transition-distribution distance in pitch and duration.The A-B-A' reading is used as a reference, while the result is framed as evidence consistent with boundary sensitivity.
  • Conclusion: The rhythmic trajectory is more selective than the pitch trajectory at the tested boundaries.The conclusion summarizes the cross-dimensional result without treating either curve as a unique segmenter.
  • Conclusion: Re-synthesis demonstrates controllable departure from the source in both pitch and duration as the window length changes.This extends the representation from analysis to generation.
  • Scope: The evidence is a first-case, multiscale observation consistent with boundary sensitivity, not a claim of unique automatic segmentation or corpus-level generality.The scope limitation is explicit in the conclusion.
Loading 2608.24660v1…