MMAE — ICASSP 2027 Audio Editing Challenge

Audio Editing Challenge

ICASSP 2027

News
  • 2026-10-03
    The registration deadline has been extended from October 1 to October 8, 2026. The Agent Track model release cutoff has been extended from October 1 to November 1, 2026. The specific versions and weights of all model components must have been publicly released before November 1, 2026.
  • 2026-09-11
    The baselines for both the Single Model Track and the Agent Track are now available! Check out the code and usage instructions on GitHub.
  • 2026-09-01
    Registration for teams is open now! Register early to receive the latest updates.
  • 2026-08-18
    Website goes live! 🚀🚀🚀

Introduction

Imagine editing audio as naturally as editing text: remove audience applause while leaving the speaker and room acoustics untouched; replace a song lyric while preserving the singer’s voice and accompaniment; or refine a recording through several dependent rounds of instructions. Recent generative audio models are beginning to make such interactions possible, moving audio production from specialized software pipelines toward an intelligent, natural-language interface.

Existing methods span domain-specific speech, music, and sound editing models, general-purpose systems that unify several editing operations and audio modalities, and planner-guided systems that decompose complex instructions into executable steps [1-13].

Current systems can often complete a simple edit in one domain, yet they remain unreliable when an instruction requires precise timing, multiple operations, mixed speech-music-sound content, multi-step reasoning, or multi-round interaction. Even when the requested change is successful, a model may alter speaker identity, musical structure, background ambience, or overall audio quality. On MMAE, existing systems achieve below 5% Exact Match Rate, revealing a wide gap between producing a plausible result and executing an edit completely and faithfully [14].

Challenge Goals

The Audio Editing Challenge at ICASSP 2027 studies how to build general-purpose audio editors that can accurately perceive and understand source audio, reason about potentially complex user instructions, and generate the requested modification while preserving everything that should remain unchanged.

The challenge features two complementary tracks:

  1. Single Model Track: Build one end-to-end model that directly transforms the input audio according to the instruction. Tasks may use any MMAE audio modality, while task complexity is restricted to the MMAE single category. This track studies intrinsic, unified audio editing capability without external models or tools.
  2. Agent Track: Build an autonomous editing system that may plan, invoke multiple locally deployed models or signal-processing tools, inspect intermediate results, and revise its output. Tasks may use any MMAE audio modality and any MMAE complexity category. This track studies orchestration, tool use, and self-correction.
Representative MMAE audio editing examples

Representative examples from the MMAE benchmark, illustrating diverse audio modalities, task complexity, editing granularity and operations, and rubric-based evaluation.

Timeline

Event Date
Registration Opens and Challenge Guidelines Released September 1, 2026
Challenge Begins October 1, 2026
Registration Deadline (Extended) October 8, 2026
Challenge Test Set and Submission SDK Released; Leaderboard Opens November 10, 2026
Final Submission Deadline and Leaderboard Freeze November 25, 2026
Evaluation and Reproducibility Check Completed December 7, 2026
Final Rankings and Invited Teams Announced December 8, 2026
Invited Two-Page ICASSP Papers Due January 7, 2027

Note: Registration and submission deadlines are at 11:59 PM on the respective day in U.S. Pacific Time. For the Agent Track, model versions and weights must have been publicly released before November 1, 2026. This tentative timeline is subject to change in accordance with the ICASSP 2027 conference schedule.

Challenge Tracks

Track 1: Single Model Track

Participants build a single, end-to-end audio editing model that consumes one or more input recordings and a natural-language instruction and directly produces the edited audio. All learned components used to understand and execute the edit must belong to one model, without delegating to separately trained perception models, planners, editors, or external services.

This track covers any MMAE audio modality, including speech, music, sound, and their mixtures. Tasks are restricted to the MMAE single complexity category.

Learn more about Track 1

Track 2: Agent Track

Participants build an autonomous audio editing agent that may orchestrate multiple open-source models and signal-processing tools, including ASR, captioning, source separation, acoustic analysis, planning, and iterative editing. An agent may maintain structured memory, inspect intermediate audio, and revise its output.

This track covers any MMAE audio modality and any MMAE complexity category, including single and complex multi-part, multi-instruction, multi-audio, multi-round, and multi-hop tasks.

Learn more about Track 2

Baselines

The challenge provides open-source baselines for both the Single Model Track and the Agent Track as reproducible starting points and experimental references, helping participants run the complete workflow.

  • Single Model Track: AuK-based end-to-end audio editing. This baseline uses the AuK base model [18] to generate edited audio directly from the input audio and a natural-language instruction. To preserve the single-model setting, Prompt Enhancer is disabled, and the output audio duration matches the input duration. It provides a reference for exploring instruction understanding and audio editing within a single model.
  • Agent Track: LLM-orchestrated audio tools. This baseline uses DeepSeek-V4-Flash as the default router to select tools according to the editing instruction. Digital signal processing (DSP) tools handle speed, volume, and pitch adjustments; SAM-Audio-Large [19] handles source separation; and AuK with Prompt Enhancer enabled [18] handles generative audio editing.

Code: Audio Editing Challenge Baselines

Benchmark and Evaluation Protocol

Benchmark

Before November 10, 2026, participants may use the publicly available MMAE test set, comprising 2,000 examples and 17,741 atomic rubrics, to develop and evaluate their models and agent systems. The dataset is available on Hugging Face.

On November 10, 2026, the organizers will release the previously unreleased challenge test set and the submission software development kit (SDK) for uploading results. The leaderboard will open for submissions on the same date. The Single Model Track and the Agent Track will each use 500 previously unreleased test examples, and the two tracks will be ranked independently. These examples are constructed through the same MMAE pipeline and manually annotated and verified. Track 1 will be evaluated on examples from any audio modality in the MMAE single complexity category; the Agent Track will be evaluated across any audio modality and any complexity category. Test inputs and editing instructions will be provided for inference, while the rubrics remain private until the official results are finalized. The complete final set and its rubrics will be released after the competition.

Submission Format

For each test item, participants submit the edited audio file and a JSONL record that maps the sample ID to its relative audio path:

{"id": "<sample_id>", "audio_path": "audio/<sample_id>.wav"}

The audio files and JSONL manifest are packaged together and uploaded to the challenge leaderboard. The two tracks are ranked independently.

Evaluation Metrics

Each rubric is evaluated three times by a frozen Qwen3-Omni model serving as the judge, with shuffled answer choices and majority voting [17]. The leaderboard reports:

  1. Instruction Following Rate (IFR): the average score over rubrics that verify whether the requested edits were correctly executed.
  2. Consistency Rate (CR): the average score over rubrics that verify whether unrelated audio content and quality were preserved.
  3. Exact Match Rate (EMR): the proportion of samples for which all instruction-following and consistency rubrics are satisfied.

Systems are ranked primarily by Overall EMR, with ties broken first by Overall IFR and then by Overall CR.

Registration and Leaderboard

The registration deadline has been extended from October 1 to October 8, 2026. To participate, please complete the registration form. Register early to receive the latest challenge updates.

Before submissions open on November 10, 2026, the organizers will send a form to collect and confirm each team’s final member list, including supervisors and team leaders. Except in special circumstances, changes to team membership will not be permitted once submissions open. Each person may participate in both tracks, but may be listed on only one team per track.

Learn more about Leaderboard

Paper Submission

The top three teams in each track will be invited to submit a two-page ICASSP 2027 paper and present their work in the Grand Challenge session.

Sponsorship

Tencent

The challenge prize pool is sponsored by Tencent, with a total of USD 7,000. The top three teams in the Single Model Track and the Agent Track will be awarded separately, with the following prizes in each track:

  • First Prize (1st place): USD 2,000
  • Second Prize (2nd place): USD 1,000
  • Third Prize (3rd place): USD 500

The organizers thank Tencent for supporting this challenge and research in audio editing.

Contact

We have a Slack workspace and a WeChat group for real-time communication. For private questions, or if an invitation link or QR code has expired, please contact Zhikang Niu or Wenming Tu.

Slack workspace QR code
Slack Workspace
WeChat group QR code
WeChat Group
Scan to join

Organizers

Zhikang Niu
Shanghai Jiao Tong University
Wenming Tu
Shanghai Jiao Tong University
Beijing Institute for General Artificial Intelligence
Ziyang Ma
Shanghai Jiao Tong University
Nanyang Technological University
Ruiyang Xu
Shanghai Jiao Tong University
Hankun Wang
Shanghai Jiao Tong University
Bohan Li
Shanghai Jiao Tong University
Ruiqi Yan
Shanghai Jiao Tong University
Zilong Zheng
Beijing Institute for General Artificial Intelligence
Chunxiang Jin
Chunxiang Jin
Inclusion AI, Ant Group
Pengcheng Zhu
Ant Group
Hung-yi Lee
National Taiwan University
Jinyu Li
Microsoft Corporation
Carlos Busso
Carnegie Mellon University
Kai Yu
Shanghai Jiao Tong University
Eng Siong Chng
Nanyang Technological University
Xie Chen
Shanghai Jiao Tong University

References

  1. Peng, Puyuan, et al. "VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild." Proc. ACL (2024).
  2. Zhang, Yixiao, et al. "MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models." arXiv:2402.06178 (2024).
  3. Wang, Yuancheng, et al. "AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models." Proc. NeurIPS (2023).
  4. Tao, Ye, et al. "MMEdit: A Unified Framework for Multi-Type Audio Editing via Audio Language Model." arXiv:2512.20339 (2025).
  5. Tian, Zeyue, et al. "Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing." Proc. SIGGRAPH (2026).
  6. Lan, Zitong, et al. "Guiding Audio Editing with Audio Language Model." Proc. NeurIPS (2025).
  7. Qiang, Chunyu, et al. "UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions." Proc. ACL (2026).
  8. Gong, Junmin, et al. "ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation." arXiv:2602.00744 (2026).
  9. Li, Zhaoqing, et al. "UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion." arXiv:2605.31530 (2026).
  10. Chen, Junyang, et al. "CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models." arXiv:2601.05329 (2026).
  11. Zhang, Dong, et al. "MiMo-Audio: Audio Language Models are Few-Shot Learners." arXiv:2512.23808 (2025).
  12. Chen, William, et al. "AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing." arXiv:2602.17097 (2026).
  13. Wang, Hankun, et al. "dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model." arXiv:2608.02673 (2026).
  14. Ma, Ziyang, et al. "MMAE: A Massive Multitask Audio Editing Benchmark." arXiv:2606.07229 (2026).
  15. Yan, Chao, et al. "Step-Audio-EditX Technical Report." arXiv:2511.03601 (2025).
  16. Yan, Canxiang, et al. "Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation." arXiv:2511.05516 (2025).
  17. Xu, Jin, et al. "Qwen3-Omni Technical Report." arXiv:2509.17765 (2025).
  18. Ma, Ziyang, et al. "AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing." arXiv:2609.08936 (2026).
  19. Shi, Bowen, et al. "SAM Audio: Segment Anything in Audio." arXiv:2512.18099 (2025).

Follow us on GitHub for updates: @Audio-Editing-Challenge