Source-linked AI summary

A Functional Taxonomy of Music Generation Systems

Dorien Herremans, Ching-Hua Chuan, Elaine Chew

arXiv:1812.04186v1cs.SDcs.LGeess.AS

TL;DR

Automatic music generation remains difficult to compare systematically because researchers still debate the tasks machines solve and the quality of their outcomes. This paper proposes a functional taxonomy organized by system purpose, surveys existing systems through that lens, and identifies uncharted areas and challenges for future work. The overview highlights long-term structure, higher-level content, reasoning, training-data efficiency, and transparency as continuing challenges.

  • Problem

    Previous surveys do not provide a systematic comparison of music-generation tasks, modeling, system relationships, objectives, and evaluation.

  • Method

    The paper proposes a functional taxonomy and surveys systems according to their purposes, functions, design aspects, and relationships rather than primarily by algorithmic technique.

  • Results

    The taxonomy provides a functional overview that identifies uncharted areas and challenges, including long-term structure, higher-level content, reasoning, training-data needs, and transparent systems.

  • Takeaways & Limitations

    A function- and design-centered perspective clarifies the frontiers of automatic music generation and sets the stage for new breakthroughs.

  • Takeaways & Limitations

    Evaluating generated music continuously with human ratings is impractical because it requires excessive time and can cause listener fatigue.

Abstract

from arXiv · show

Digital advances have transformed the face of automatic music generation since its beginnings at the dawn of computing. Despite the many breakthroughs, issues such as the musical tasks targeted by different machines and the degree to which they succeed remain open questions. We present a functional taxonomy for music generation systems with reference to existing systems. The taxonomy organizes systems according to the purposes for which they were designed. It also reveals the inter-relatedness amongst the systems. This design-centered approach contrasts with predominant methods-based surveys and facilitates the identification of grand challenges to set the stage for new breakthroughs.

1. INTRODUCTION

Automatic music generation remains an ill-defined research problem because the tasks, objectives, modeling choices, connections, and evaluation of systems are not systematically compared. The paper addresses this gap with a function- and design-centered taxonomy that organizes systems by purpose and relates their compositional roles.

  • Research gap: Previous surveys do not systematically compare the musical tasks machines perform, how those tasks are modeled, how systems connect, or how outcomes are evaluated.The paper identifies these as outstanding questions for the field.
  • Approach: The taxonomy classifies automatic music generation systems by their functions and design purposes rather than by algorithmic techniques.This perspective contrasts with surveys organized around methods such as Markov models, genetic algorithms, rules, or neural networks.
  • Taxonomy structure: The concept map centers on composition and note, with melody, harmony, rhythm, and timbre as four essential compositional elements between them.Notes are characterized by properties including pitch, duration, onset time, and instrumentation.
  • Taxonomy structure: Systems may address multiple functional aspects simultaneously or target one aspect while treating others as fixed or user-provided.The objective selected for a system affects its problem definition and prospective solution techniques.
  • High-level functions: Interactive composing uses user input in an online music-generation process that may be real-time or not.Examples include learning a player’s style while improvising and using feedback as critique or generation parameters.
  • High-level functions: Narrative, interaction, and difficulty extend the taxonomy above composition, covering listener-perceived story or emotion, user participation, and instrument playability.Long-term or hierarchical structure supports these higher-level goals.

2. A FUNCTIONAL INDEX OF MUSIC GENERATION SYSTEMS

The functional index surveys selected music-generation systems by the musical functions they address and the techniques used within those functions. It emphasizes each system’s most prominent contribution while recognizing that systems and functional aspects can overlap.

  • Functional scope: The index covers melody, harmony, rhythm, timbre, interaction, narrative, difficulty, and long-term structure.These functional aspects organize the survey of example systems.
  • Functional overlap: Functional aspects can be conflated, because a system discussed under one aspect may also address others, such as rhythm within melody.This overlap means the index’s categories are not mutually exclusive.
  • Index organization: Table I classifies selected systems by their main technique and lists each with its most prominent functional aspect.Only systems with a clear contribution are included, and systems may belong to more than one category.
  • Technique categories: Rule-, constraint-satisfaction-, and grammar-based systems form one technique category in the functional overview.The table also includes neural, evolutionary, local-search, and other optimization categories.
  • Technique categories: Neural-network systems are represented across functional areas including harmony, melody, and interaction.The listed neural approaches include restricted Boltzmann machines and LSTM models.
  • Technique categories: Evolutionary and population-based optimization systems span melody, harmony, rhythm, interaction, difficulty, and timbre.The index associates these functions with the cited systems in the overview.

2.1. Melody

Melody-generation systems range from stochastic and rule-based methods to optimization and neural models, increasingly addressing style, repetition, structure, and accompaniment constraints. The section highlights a persistent tension between similarity to source music and originality.

  • Melodic generation: Melody systems commonly target similarity to a chosen style or corpus, using extracted features, statistical models, or music-theoretic rules.Examples include Western tonal music, free jazz, nursery rhymes, hymns, and bagana music.
  • Melodic generation: Higher-order Markov models tend to generate more repetitive melodies, whereas lower-order models produce more randomness.
  • Melodic generation: Automatic music generation must balance resemblance to a target style with sufficient novelty to avoid excessive repetition and plagiarism.MaxOrder limits subsequence length in higher-order Markov generation to curb repeated source material.
  • Melodic generation: Long-term melodic structure requires explicit patterns, motives, variations, or structural constraints beyond random-walk or Gibbs-sampling generation.Several systems use concatenated patterns, integer programming, optimization, or tension and repetition constraints to address structure.
  • Melodic generation: IDyOM combines long- and short-term Markov models so generated melodies reflect both corpus statistics and structure emerging within the current piece.Its short-term model increases self-similarity and can support a precursor to musical form.
  • Melodic generation: Impro-Visor listeners matched 95% of Clifford Brown-style solos, 90% of Miles Davis-style solos, and 85% of Freddie Hubbard-style solos to the original performer.The system balances similarity and novelty but does not capture long-term structure.

2.2. Harmony

Harmony-generation systems model chords, cadences, counterpoint, chorale harmonization, and chord sequences according to stylistic and voice-leading constraints. Approaches include explicit rules, optimization-based evaluation, and machine learning.

  • Harmony: Harmony systems generate chords and cadences according to a target style, with quality often judged by harmonic similarity and voice-leading rules.Popular-music chord progressions also depend on vertical melody–harmony relations and horizontal chord-transition patterns.
  • Counterpoint: Counterpoint composes one or more independent voices against a given cantus firmus under a strict set of melodic and harmonic rules.Systems commonly handle two to four voices, while four-part counterpoint is grouped with chorale harmonization.
  • Counterpoint: Counterpoint systems emulate style through explicit rules, rule-based optimization, or machine learning.The approaches respectively generate with known rules, optimize rule adherence, or learn stylistic regularities from examples.
  • Counterpoint: Eighteen melodic and fifteen harmonic rules based on Fux’s theory generate a cantus firmus and first-species counterpoint in Herremans and Sörensen’s system.The rules were implemented in an objective function optimized with variable neighborhood search.
  • Chorale harmonization: Chorale harmonization most commonly generates three voices designed to harmonize a given soprano melody.Bach-in-a-Box instead accepts a user-created melody in any of four voices and generates the other three notes for each chord.

2.3. Rhythm

Rhythm-generation research is less common than melody- and harmony-focused work and often treats rhythm as an attribute of note events. Existing systems use interactive, evolutionary, Markov, and audio-analysis approaches.

  • Rhythm: Fewer systems solely generate rhythm than systems focusing on melody or harmony, although similar modeling approaches are used.Rhythm is often given or embedded as an attribute of note events.
  • Rhythm: CONGA evolves rhythmic fragments with genetic algorithms and determines their concatenation through genetic programming, using user feedback as fitness.Participants used the system to produce rhythmic progressions sounding like rock ’n’ roll.
  • Rhythm: Ariza’s genetic algorithm measures distance from a user-provided fit-rhythm through five distance measures, but the paper reports no evaluation results.
  • Rhythm: Markov-based systems learn rhythmic patterns from users, drummers, or audio examples and reproduce likely sequences with stylistic or structural consistency.Representations may include duration, velocity, instrument, onset timing, and symbolic patterns derived from audio.
  • Rhythm: Interactive rhythm generation can analyze a human drummer in real time and select among imitation, transformation, detection, and accompaniment modes.Haile uses six interaction modes based on perceptual analysis of the human player.

2.4. Timbre

Timbre generation addresses tone color in orchestration and sound synthesis, where systems search or retrieve combinations of instrument sounds matching a target timbre. Later work also incorporates audio-segment analysis and final-mix quality.

  • Timbre: Timbre distinguishes voices and instruments and affects perception of an orchestra’s composite sound.
  • Timbre: Target-timbre orchestration is commonly modeled as a combinatorial search through instrument sound samples for a perceptually similar combination.
  • Timbre: Early systems measured timbral closeness primarily through frequency-spectrum similarity and retrieved or combined instruments from feature-described databases.Features included pitch range, dynamic levels, clef, and transposition information.
  • Timbre: Carpentier et al. model orchestration as a constrained multi-objective search over sound attributes and features.The objective is to find sound combinations similar to a target timbre.
  • Timbre: Collins applies machine learning to electroacoustic composition by analyzing, combining, and modifying audio segments while accounting for final-mix quality.Modifications include delays, filters, and time stretching.

2.5. Interaction

Interactive music-generation systems listen and respond to performers or other inputs in real time, spanning structured and free improvisation. Systems differ in how they learn, anticipate, accept feedback, and represent musical material.

  • Interactive systems create music while listening to a human performer, combining anticipation, improvisation, and two-way communication.
  • Structured improvisation: Structured improvisation systems generate within predefined musical settings such as chord progressions or fixed-length jazz solos.GenJam evolves a player’s last four bars with a genetic algorithm, while BoB models pitch class, interval, and melodic direction probabilistically.
  • Structured improvisation: Other structured systems combine rules, neural networks, and reinforcement learning to trade phrases with musicians, while retaining stochastic variation.CHIME supports interactive jazz generation but its authors note limitations in its hard-coded rules and handling of out-of-chord changes.
  • Free improvisation: Free-improvisation systems learn and generate music concurrently without a fixed predefined structure, often adapting to a performer’s style.The Continuator uses an adapted Lempel-Ziv parsing algorithm that handles rhythm, beat, harmony, and imprecision.
  • Other interactive systems: Factor-oracle systems such as OMax encode and generate a player’s style concurrently, with later systems extending the approach to polyphony, audio, rhythm, and harmony.ImproteK adds rhythmic and harmonic generation and can revise an offline scenario during performance.
  • Performer feedback: Performer feedback can expose musical context and give users direct control over learning, memory, and multiple generated streams.Mimi displays recent past and future music, while Mimi4x lets users coordinate four interacting instances.

2.6. Narrative

Narrative music-generation systems organize musical material around stories, scenes, emotions, tension, game states, or recurring motifs. The surveyed approaches range from cross-fading and motif embedding to systems that generate music in response to media content.

  • Narrative music uses representational, organizational, and discursive cues to deliver story information through music.The survey considers tension profiles, fragment blending, leitmotifs, and film music as narrative forms.
  • Narrative cues create musical structure through emotional variation, tension profiles, leitmotifs, repeated patterns, and synchronization with media.These cues can produce similarity within a piece and between musical emotions and simultaneous video or games.
  • Game and video music aims to match the content or emotional content of a scene or narrative.Examples include music that changes with game scenarios or a player’s style of play.
  • Tension: Tension-based systems model or target perceptual tension using musical and audio features, profiles, motifs, or tonal representations.MorpheuS constrains detected patterns to fit a user-provided or template-derived tension profile.
  • Blending: Cross-fading commonly transitions between game-state audio files, but smooth blending requires harmonically and rhythmically similar fragments.Restricting rhythms and harmonies improves blending while reducing musical variation and expressive capacity.
  • Leitmotifs: Leitmotif systems associate distinctive musical fragments with game characters or elements and embed them in forms expressing different tension and regularity.
  • Film and video music: Video-background systems can generate harmony, melody, rhythm, and sound effects for scenes using mood, intensity, key, and movement characteristics.

2.7. Difficulty

Music-generation systems can target the difficulty of performing a piece, not only its musical features. This supports composing music for a specified skill level or evaluating generated pieces by playability.

  • Playing difficulty is the skill level required for a musician to perform a piece, while ergonomic goals concern ease of playing.
  • Generating music with a target difficulty requires considering playability of note combinations on a particular instrument.The survey identifies melody, harmony, rhythm, and timbre as common composition targets whose ergonomic implications may otherwise be overlooked.
  • Goals: Difficulty can be defined relative to a corpus’s playability or specified directly as a target level for an instrument.The authors suggest explicitly measuring difficulty rather than assuming that corpus-trained models preserve playability.
  • Examples: Genetic algorithms have generated playable guitar music by minimizing hand and finger movements.
  • Examples: Piano-piece difficulty can be measured from characteristics including harmony, fingering, polyphony, and rhythmic irregularity.The survey notes that automatic fingering systems could extend such difficulty modeling to piano and string instruments.

3. FUTURE CHALLENGES

The survey identifies future challenges beyond generating isolated musical features: modeling higher-level concepts, improving data efficiency and playability, rendering realistic audio, and enabling objective system comparison. It argues that broader intelligence and evaluation infrastructure are needed for practical adoption.

  • State-of-the-art systems generate well-defined musical aspects effectively, but everyday adoption remains an open challenge.
  • Long-term structure remains important, with recurring themes, motifs, patterns, form, cadence, and pitch contour identified as targets for further modeling.The survey specifically calls for further investigation of RNN and LSTM abilities to capture long-term structure.
  • Higher-level narrative intelligence remains a major challenge because systems still need to connect generated music with emotion, games, film, and video.The paper identifies real-time game music and background music as potential applications of this direction.
  • Machine-learning approaches to modeling musical emotion usually require large amounts of data, motivating work on systems that need less data.
  • Playing difficulty is often neglected but could support skill-tailored composition and serve as an evaluation measure for generated music.
  • Realistic rendering of generated MIDI is important for making automatically generated music attractive and usable in real applications.
  • Objective comparison is difficult because systems differ in inputs, outputs, training styles, available audio, and listener familiarity.The authors established a public repository collecting system details, outputs, and manual corrections to facilitate comparison.

4. CONCLUSIONS

The paper uses a functional taxonomy to survey what music-generation systems can and cannot do, clarifying the field’s frontiers and identifying challenges. It also highlights opportunities for advancement and supports evaluation through an online repository of generated music.

  • The functional taxonomy and survey clarify the frontiers of automatic music generation by focusing on systems’ capabilities rather than algorithmic techniques.This approach identifies uncharted areas and challenges for automatic music composition.
  • Current challenges include generating long-term structure, capturing emotion and tension, reducing training-data needs through innate reasoning, and promoting transparent and objective evaluation methods.
  • The functional overview identifies opportunities for further advancement toward applications in artistic innovation and adaptive, copyright-free music for games and videos.
  • The authors established an online repository of computer-generated music to stimulate visibility and evaluation of current systems.
Loading 1812.04186v1…