Source-linked AI summary

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hanlin Zhang, Daxin Tan, Dehua Tao, Xiao Chen, Haochen Tan, Linqi Song

arXiv:2606.01804v2eess.AScs.SD

TL;DR

Instruction-guided speech editing lacks a unified evaluation framework for modifying target attributes while preserving unrelated speech characteristics. SpeechEditBench addresses this with a bilingual multi-attribute benchmark and anchor-based metrics, revealing fragmented model capabilities, preservation bottlenecks, and persistent difficulty with compositional editing.

  • Problem

    Speech editing lacks a unified benchmark with comparable tasks and metrics that jointly assess target modification and preservation of unrelated speech characteristics.

  • Method

    SpeechEditBench evaluates seven atomic and compositional editing tasks using anchor-based metrics for target success, preservation success, and joint success.

  • Results

    Evaluations reveal fragmented capabilities, preservation as a central bottleneck, and low joint success on compositional editing even for the strongest open-source model.

  • Takeaways & Limitations

    SpeechEditBench provides a unified testbed for diagnosing speech-editing capabilities and guiding improvements in controllability and preservation.

  • Takeaways & Limitations

    SpeechEditBench currently covers only English and Chinese, excluding low-resource languages and broader language families.

Abstract

from arXiv · show

Instruction-guided speech editing requires a model to modify specified speech attributes while preserving unrelated characteristics. Despite rapid progress in Speech Large Language Models (Speech LLMs), systematic evaluation of this capability remains challenging, as existing benchmarks are fragmented across isolated editing tasks. To bridge this gap, we introduce SpeechEditBench, a bilingual multi-attribute benchmark for instruction-guided speech editing. SpeechEditBench encompasses seven atomic editing tasks, as well as compositional editing tasks that integrate multiple operations within a single instruction. We propose an anchor-based evaluation protocol that separately assesses the edit success of target attributes and the preservation of untargeted attributes, leading to three metrics: target success, preservation success, and joint success. Using this benchmark, we evaluate mainstream Speech LLMs and specialized speech editing systems. The results reveal three key findings: (1) no single model performs well across all editing dimensions; (2) closed-source Speech LLMs generally outperform open-source models; (3) compositional editing remains highly challenging, with even the most advanced models struggling to achieve high joint success. SpeechEditBench provides a rigorous diagnostic framework to identify bottlenecks in Speech LLMs, thereby facilitating the development of next-generation Speech LLMs with more robust and precise instruction-guided editing capabilities. Data and code are avaialble at https://github.com/daxintan-cuhk/SpeechEditBench .

1 Introduction

Instruction-guided speech editing is under-explored because it must modify targeted attributes while preserving unrelated speech characteristics, yet lacks a unified, fine-grained evaluation framework. Existing evaluations use incomparable task-specific metrics, rigid waveform matching, and rarely assess editing effectiveness alongside source preservation.

  • Motivation: Speech editing requires targeted attribute modification while preserving unrelated speech characteristics, alongside instruction comprehension, attribute localization, and precise semantic and acoustic manipulation.This creates dual constraints unlike conventional speech tasks with standalone objectives.
  • Problem: The field lacks a dedicated unified benchmark with consistent metrics, comprehensive tasks, and diagnosis of model strengths and weaknesses.Although audio editing benchmarks exist, they do not provide unified task definitions or consistent evaluation.
  • Problem: Existing evaluation metrics are task-specific and incomparable across studies, limiting meaningful comparison of speech-editing systems.This is identified as the first inherent flaw in current evaluation paradigms.
  • Problem: Rigid waveform matching cannot accommodate the one-to-many nature of valid speech edits.Different waveforms may represent valid edits, making exact waveform comparison unsuitable.
  • Contribution: Existing evaluations rarely consider editing effectiveness and source preservation simultaneously, motivating a more expansive and fine-grained evaluation framework.The benchmark is introduced to bridge this evaluation gap.

2 Related Work

Related work spans instruction-guided editing benchmarks in images and audio, alongside speech-editing systems evolving from spectrogram inpainting toward discrete-token, multilingual, and broader SpeechLLM-based editing.

  • Image and multimodal editing benchmarks: Image-editing benchmarks evaluate instruction execution and content preservation across objects, attributes, scenes, edit types, and multimodal interactions.Examples include EditBench, EditVal, Complex-Edit, EditInspector, GIE-Bench, and VIBE.
  • SpeechLLM benchmarks: SpeechLLM benchmarks increasingly probe robustness, instruction following, generative comprehension, and conversational capabilities across diverse audio modalities.VoiceBench covers speaker characteristics, noise, and disfluencies, while AIR-Bench and AudioBench span speech, natural sounds, and music.
  • Speech editing systems: Early text-based speech editing systems framed editing as spectrogram inpainting, emphasizing boundary smoothness and local prosody.EditSpeech and FluentEditor are representative systems in this paradigm.
  • Speech editing systems: Discrete-speech-token models shifted speech editing toward end-to-end instruction-driven editing, introducing token infilling and multilingual capabilities.VoiceCraft introduced token infilling and the RealEdit dataset; VoiceCraft-X extended multilingual editing to 11 languages.
  • Speech editing systems: Recent work also includes dedicated architectures and SpeechLLMs that unify speech editing with broader understanding and generation.Examples include CosyEdit, a 400M-parameter model fine-tuned from CosyVoice, MAVE’s cross-attentive Mamba architecture, and Ming-UniAudio.

3 Benchmark Design

SpeechEditBench unifies bilingual atomic and compositional speech-editing tasks under anchor-based evaluation that measures both target-attribute editing and preservation of untargeted content. Its design spans seven editing dimensions, cross-category compositions, and task-adaptive success criteria.

  • Design principles: The benchmark uses unified task formulation, anchor-based evaluation, and dual-constraint metrics balancing editing effectiveness with preservation fidelity.Each sample pairs source speech with a natural-language instruction, and speaker editing additionally provides a reference speech for target timbre.
  • Task structure: SpeechEditBench covers English and Chinese, with atomic instructions targeting one attribute and compositional instructions combining two or three operations.Compositional editing tests multi-constraint satisfaction and joint instruction fulfillment.
  • Atomic tasks: The seven atomic tasks edit content, speaker identity, emotion, style, prosody, paralinguistic events, or acoustic conditions.Acoustic editing includes enhancement and environment transfer, while emotion editing includes a challenging setting requiring semantic polarity override.
  • Compositional split: The compositional split contains 320 two-component and 80 three-component cross-category samples balanced across English and Chinese.The split supports evaluation of per-component success and joint instruction fulfillment.
  • Evaluation protocol: Evaluation reports target success, preservation success, and joint success as averaged binary indicators, with task-adaptive thresholds and a 10% WER/CER preservation limit for non-content edits.For compositional samples, component-wise success, all-component success, and joint success are additionally reported.

4 Evaluated Models

The evaluation covers two model categories: general-purpose SpeechLLMs for cross-task editing and specialized speech editing systems whose interfaces align with benchmark tasks. Six open-source and two closed-source SpeechLLMs are assessed, alongside task-specific reference systems for individual editing attributes.

  • Model categories: The benchmark evaluates two categories: SpeechLLMs and specialized speech editing systems.SpeechLLMs target general cross-task editing capability, while specialized systems serve as reference systems when their native interfaces align with a SpeechEditBench task.
  • SpeechLLMs: Six open-source and two closed-source SpeechLLMs are evaluated for general cross-task editing.The open-source systems are Ming-UniAudio, Step-Audio-EditX, Qwen3-Omni, Kimi-Audio, Mimo-Audio-Base, and Mimo-Audio-Instruction; the closed-source systems are Gemini-Live and GPT-Realtime.
  • SpeechLLMs: Speaker editing is not tested for the evaluated SpeechLLMs because they do not accept both source and reference audio as inputs.This limitation applies to the six open-source and two closed-source SpeechLLM systems listed in the evaluation.
  • Specialized speech editing systems: Specialized systems cover content, speaker, emotion/style/prosody, paralinguistic, and enhancement-related editing tasks.VoiceCraft-X handles content; Seed-VC handles speaker; VoxCPM2 handles emotion, style, and prosody; Chatterbox and AudioSep handle paralinguistic editing; Deep-FilterNet and a digital signal processing algorithm provide additional references.

5 Results and Discussion

Results show fragmented capability across editing dimensions: closed-source SpeechLLMs generally preserve untargeted attributes better, while no model dominates every task. Compositional editing remains especially difficult, and specialized systems provide strong but non-uniform task-specific references.

  • Overall performance: Table 2 exposes a target-preservation gap: models may achieve requested edits while failing to preserve non-target attributes.Joint success combines target achievement with applicable preservation, whereas target success measures edit completion.
  • Open-source models: Open-source performance is fragmented: Ming-UniAudio leads content at 76.46% joint, while Step-Audio-EditX leads emotion, style, and paralinguistic editing.Step-Audio-EditX achieves 7.71%, 49.67%, and 31.25% joint success on emotion, style, and paralinguistic editing, respectively.
  • Closed-source models: GPT-Realtime leads content, style, and paralinguistic joint success at 96.67%, 68.67%, and 47.00%, while Gemini-Live leads emotion and prosody at 27.79% and 65.17%.Qwen3-Omni slightly exceeds both closed-source systems on acoustic joint success at 37.80%, so the closed-source advantage is not uniform.
  • Preservation trade-offs: Preservation differentiates system families: GPT-Realtime reaches 81.50% target and 47.00% joint success on paralinguistic editing, versus Step-Audio-EditX at 61.75% and 31.25%.Mimo-Audio-Instruction similarly achieves high target success on several tasks but near-zero joint success because it fails to preserve non-target attributes.
  • Specialized systems: Specialized systems remain strong references: VoiceCraft-X reaches 84.00% English content joint success, and Seed-VC reaches 80.50% speaker joint success under reference-based evaluation.Their strengths are task-specific and do not uniformly surpass SpeechLLMs, especially across expressive tasks.
  • Compositional editing: Compositional editing remains difficult: Gemini-Live and GPT-Realtime achieve 21.50% and 20.00% joint success on two-component combinations, while Qwen3-Omni reaches 10.00%.Component success is much higher than all-component or joint success, and nearly all models obtain 0.00% joint success on the harder setting.

6 Conclusion

SpeechEditBench is a bilingual benchmark with anchor-based metrics for diagnosing instruction-guided speech editing. Evaluation reveals fragmented model capabilities, preservation as a central bottleneck, and unresolved compositional editing.

  • Benchmark and evaluation: SpeechEditBench comprises 4,700 bilingual samples spanning seven atomic tasks and compositional editing, with separate target, preservation, and joint success metrics.Its anchor-based protocol evaluates both whether requested attributes are edited and whether untargeted attributes are preserved.
  • Key findings: No evaluated model performs reliably across all editing dimensions, revealing severe capability fragmentation.The conclusion reports evaluations of eight Speech LLMs and task-specific systems.
  • Key findings: Preservation of untargeted attributes remains the central bottleneck, as successful edits often corrupt lexical content.This finding specifically concerns preserving characteristics that the instruction does not target.
  • Key findings: Compositional editing remains far from solved, with even the strongest open-source model achieving low joint success.The benchmark therefore exposes difficulties in performing multiple editing operations while preserving unrelated attributes.
  • Implications: SpeechEditBench offers a unified testbed for diagnosing these gaps and guiding future speech foundation models toward stronger controllability and preservation.The dataset and evaluation code are planned for release upon acceptance.

7 Limitations · A Construction and Evaluation Details · A.1 Dataset Composition

SpeechEditBench is limited to English and Chinese, automatic evaluation, single-turn editing, and seven atomic task families with compositions. Its appendix details construction and reports 4,700 filtered, balanced, deduplicated samples spanning five source-data categories.

  • 7 Limitations: SpeechEditBench supports only English and Chinese, excluding low-resource languages and broader language families.The authors describe the two languages as typologically distinct but acknowledge limited language coverage.
  • 7 Limitations: Automatic WER/CER, speaker-similarity, and Gemini-based metrics cannot fully capture perceptual naturalness or subtle prosodic and paralinguistic nuances.Future work will incorporate human listening tests for finer-grained subjective validation.
  • 7 Limitations: The benchmark evaluates single-turn instruction-guided editing and excludes multiturn refinement based on conversational feedback.The paper identifies iterative editing as important for natural spoken interaction.
  • 7 Limitations: Its scope covers seven atomic editing tasks and compositions but excludes code-switching, dialectal accent conversion, and fine-grained voice-style interpolation.These omitted dimensions may matter for future specialized applications.
  • A Construction and Evaluation Details: The appendix supplements the main paper with sample distributions, construction criteria, filtering prompts, evaluation thresholds, and qualitative failure types.These details are presented as methodological additions rather than expanded main-text content.
  • A.1 Dataset Composition: 4,700 samples comprise the benchmark after filtering, balancing, and deduplication: 4,300 atomic samples and 400 compositional samples.The composition is summarized in Table 7.
  • A.1 Dataset Composition: Five source-data categories support distinct editing dimensions, including clean read speech, emotional and expressive corpora, speaker-labeled corpora, nonverbal speech, and acoustic resources.These categories provide content, prosody, acoustic, expressive, identity, paralinguistic, noise, and room-response resources.
  • A.1 Dataset Composition: Acoustic resources provide additive noises and room responses for synthesizing controlled acoustic conditions, while speaker-labeled corpora provide identity anchors.Nonverbal corpora supply paralinguistic events, and expressive corpora contribute emotion- and style-rich utterances.

A.2 Atomic Task Construction · A.3 Filtering Prompt Specifications

SpeechEditBench constructs atomic-task samples from source speech, natural-language instructions, and verifiable task-specific anchors. Its filtering prompts enforce semantic completeness and task-specific criteria for content, emotion, style, and paralinguistic events.

  • A.2 Atomic Task Construction: Each atomic-task sample contains a source speech signal, a natural-language instruction, and task-specific anchors defining verifiable targets.
  • A.3 Filtering Prompt Specifications: Filtering precedes construction for several tasks, using shared criteria and structured response fields across language-specific prompt variants.
  • A.3 Filtering Prompt Specifications: Content candidates must be complete, self-contained utterances, while replacement edits target one content word or 1–3-word phrase that can be naturally substituted.
  • A.3 Filtering Prompt Specifications: Neutral-emotion filtering requires semantic completeness and no clearly expressed emotion based on text alone, returning semantic_complete, is_neutral, score, and reason.
  • A.3 Filtering Prompt Specifications: Challenging-emotion filtering requires semantic completeness and clear textual expression of the target emotion, returning semantic_complete, is_salient, score, and reason.
  • A.3 Filtering Prompt Specifications: Both emotion-filtering modes select an utterance only when the boolean decision is true and the score is at least 70.
  • A.3 Filtering Prompt Specifications: Style annotation evaluates delivery rather than emotion or topic, distinguishes closely related delivery modes, and scores six styles independently from 0 to 4.
  • A.3 Filtering Prompt Specifications: Paralinguistic annotation scores breath, laugh, cough, and sigh independently from 0 to 3, with add operations requiring absent-or-faint events and remove operations requiring noticeable-or-prominent events.

B Compositional Task Construction · C Anchor Semantics

SpeechEditBench constructs compositional instructions by combining verified atomic samples while retaining component-level anchors for evaluation. Its anchor semantics distinguish intended edits from preserved attributes and support independent, all-component, and joint assessment.

  • B Compositional Task Construction: Each compositional sample uses one source utterance, while other verified atomic samples provide target attributes such as speaker, emotion, prosody, or acoustic environment.This preserves atomic anchors for evaluating individual components.
  • B Compositional Task Construction: The compositional split organizes edits into four attribute groups: semantic content, speaker identity, expressive delivery, and acoustic environment.
  • B Compositional Task Construction: Table 9 specifies the combinations included in the compositional split.
  • B Compositional Task Construction: Across 400 compositional samples, component counts are content 220, speaker 180, emotion 180, prosody 80, and acoustic 220.
  • B Compositional Task Construction: Prosody controls cover speed and pitch, with 20 examples each for faster, slower, higher, and lower directions.
  • C Anchor Semantics: Anchor-based evaluation specifies verifiable target and preservation requirements without requiring a unique target waveform.A target anchor defines the intended edit, whereas a preservation anchor defines what should remain unchanged.
  • C Anchor Semantics: Speaker references function as identity anchors, and acoustic references or conditions describe measurable environments rather than unique waveform targets.Compositional components are evaluated independently and then combined into all-component and joint success.

D Task Examples … E.5 Diagnostic Metrics

The evaluation protocol combines binary target, preservation, and joint success metrics with task-specific evaluators and content-preservation gates. It supplements primary criteria with multimodal judgments and naturalness, speaker, intelligibility, and quality diagnostics.

  • D Task Examples: Table 11 provides representative task instances and reports only fields relevant to each task.
  • E Evaluation Protocol Details: The protocol evaluates each sample using target success, preservation success, and joint success, abbreviated TS, PS, and JS.For content editing, edit success is primary, while exact match and WER/CER are diagnostics rather than preservation gates.
  • E.1 Success Metrics: For content editing, the edited transcript is the target, so edit success is primary and exact match plus WER/CER are diagnostic measures.
  • E.2 Content Preservation Gate: Non-content atomic tasks impose content preservation as a hard gate using ASR, language-specific transcription models, and normalized transcript comparison.English uses Whisper large-v3; Chinese uses Paraformer through FunASR, with punctuation and spacing normalized according to language.
  • E.3 Atomic Evaluators: Table 12 defines each task’s target-success criterion, while diagnostic metrics are not substitutes for primary success criteria.
  • E.4 Multimodal Judge Prompts: Emotion, style, and paralinguistic targets are judged by a temperature-0 Gemini-based multimodal evaluator returning structured JSON.The judges assess audio delivery, including emotion labels, vocal style, and speaker-produced events such as breath, laugh, cough, and sigh.
  • E.4 Multimodal Judge Prompts: Emotion judgments normalize aliases before matching targets, while style judgments score vocal delivery and paralinguistic judgments exclude background noise and microphone artifacts.
  • E.5 Diagnostic Metrics: Diagnostic evaluation reports UTMOS naturalness, WavLM-large speaker cosine similarity, DNSMOS SIG/BAK/OVRL, and PESQ/STOI when clean references exist.

F Compositional Metrics … J Licenses

The paper defines compositional metrics that distinguish component-level achievement, simultaneous satisfaction of all requested edits, and preservation-aware joint success, then documents evaluation diagnostics, AI usage, model sizes, and artifact licensing. These sections frame both how compositional performance is measured and the implementation, taxonomy, and legal context of the benchmark.

  • F Compositional Metrics: Pooled component success measures whether individual compositional sub-goals succeed under their corresponding atomic evaluators.Each component is represented by a binary success indicator a_i,c.
  • F Compositional Metrics: All-component success measures whether all requested edits are satisfied simultaneously.It is defined separately from individual component success.
  • F Compositional Metrics: Joint success additionally requires preservation of non-target content when applicable.The joint indicator combines edit satisfaction with preservation result p_i.
  • G Failure-Type Taxonomy: Table 13 provides representative diagnostic descriptions of failure types observed during evaluation.The examples are explicitly described as diagnostic rather than additional headline results.
  • H AI Usage: AI was used for grammatical checking.
  • I Model Size: Evaluated models range from Seed-VC with 98M parameters to Qwen3-Omini with 30 billion parameters, and experiments ran on Ascend.Other listed sizes include 16B, 7B, 3B, and 2B parameters; generation configurations were default.
  • J Licenses: The benchmark’s external artifacts use mixed licensing, including CC BY 4.0, Apache 2.0, CC BY-NC-SA 4.0, and CC BY-NC 4.0 with research-only restrictions for StoryTTS.The passage also identifies a custom open-data license for MagicData-RAMC.
Loading 2606.01804v2…