Audio Editing Challenge
ICASSP 2027
-
2026-10-03The registration deadline has been extended from October 1 to October 8, 2026. The Agent Track model release cutoff has been extended from October 1 to November 1, 2026. The specific versions and weights of all model components must have been publicly released before November 1, 2026.
-
2026-09-11The baselines for both the Single Model Track and the Agent Track are now available! Check out the code and usage instructions on GitHub.
-
2026-09-01Registration for teams is open now! Register early to receive the latest updates.
-
2026-08-18Website goes live! 🚀🚀🚀
Introduction
Imagine editing audio as naturally as editing text: remove audience applause while leaving the speaker and room acoustics untouched; replace a song lyric while preserving the singer’s voice and accompaniment; or refine a recording through several dependent rounds of instructions. Recent generative audio models are beginning to make such interactions possible, moving audio production from specialized software pipelines toward an intelligent, natural-language interface.
Existing methods span domain-specific speech, music, and sound editing models, general-purpose systems that unify several editing operations and audio modalities, and planner-guided systems that decompose complex instructions into executable steps [1-13].
Current systems can often complete a simple edit in one domain, yet they remain unreliable when an instruction requires precise timing, multiple operations, mixed speech-music-sound content, multi-step reasoning, or multi-round interaction. Even when the requested change is successful, a model may alter speaker identity, musical structure, background ambience, or overall audio quality. On MMAE, existing systems achieve below 5% Exact Match Rate, revealing a wide gap between producing a plausible result and executing an edit completely and faithfully [14].
Challenge Goals
The Audio Editing Challenge at ICASSP 2027 studies how to build general-purpose audio editors that can accurately perceive and understand source audio, reason about potentially complex user instructions, and generate the requested modification while preserving everything that should remain unchanged.
The challenge features two complementary tracks:
- Single Model Track: Build one end-to-end model that directly transforms the input audio according to the instruction. Tasks may use any MMAE audio modality, while task complexity is restricted to the MMAE single category. This track studies intrinsic, unified audio editing capability without external models or tools.
- Agent Track: Build an autonomous editing system that may plan, invoke multiple locally deployed models or signal-processing tools, inspect intermediate results, and revise its output. Tasks may use any MMAE audio modality and any MMAE complexity category. This track studies orchestration, tool use, and self-correction.
Representative examples from the MMAE benchmark, illustrating diverse audio modalities, task complexity, editing granularity and operations, and rubric-based evaluation.
Timeline
| Event | Date |
|---|---|
| Registration Opens and Challenge Guidelines Released | September 1, 2026 |
| Challenge Begins | October 1, 2026 |
| Registration Deadline (Extended) | October 8, 2026 |
| Challenge Test Set and Submission SDK Released; Leaderboard Opens | November 10, 2026 |
| Final Submission Deadline and Leaderboard Freeze | November 25, 2026 |
| Evaluation and Reproducibility Check Completed | December 7, 2026 |
| Final Rankings and Invited Teams Announced | December 8, 2026 |
| Invited Two-Page ICASSP Papers Due | January 7, 2027 |
Note: Registration and submission deadlines are at 11:59 PM on the respective day in U.S. Pacific Time. For the Agent Track, model versions and weights must have been publicly released before November 1, 2026. This tentative timeline is subject to change in accordance with the ICASSP 2027 conference schedule.
Challenge Tracks
Track 1: Single Model Track
Participants build a single, end-to-end audio editing model that consumes one or more input recordings and a natural-language instruction and directly produces the edited audio. All learned components used to understand and execute the edit must belong to one model, without delegating to separately trained perception models, planners, editors, or external services.
This track covers any MMAE audio modality, including speech, music, sound, and their mixtures. Tasks are restricted to the MMAE single complexity category.
Track 2: Agent Track
Participants build an autonomous audio editing agent that may orchestrate multiple open-source models and signal-processing tools, including ASR, captioning, source separation, acoustic analysis, planning, and iterative editing. An agent may maintain structured memory, inspect intermediate audio, and revise its output.
This track covers any MMAE audio modality and any MMAE complexity category, including single and complex multi-part, multi-instruction, multi-audio, multi-round, and multi-hop tasks.
Baselines
The challenge provides open-source baselines for both the Single Model Track and the Agent Track as reproducible starting points and experimental references, helping participants run the complete workflow.
- Single Model Track: AuK-based end-to-end audio editing. This baseline uses the AuK base model [18] to generate edited audio directly from the input audio and a natural-language instruction. To preserve the single-model setting, Prompt Enhancer is disabled, and the output audio duration matches the input duration. It provides a reference for exploring instruction understanding and audio editing within a single model.
- Agent Track: LLM-orchestrated audio tools. This baseline uses DeepSeek-V4-Flash as the default router to select tools according to the editing instruction. Digital signal processing (DSP) tools handle speed, volume, and pitch adjustments; SAM-Audio-Large [19] handles source separation; and AuK with Prompt Enhancer enabled [18] handles generative audio editing.
Code: Audio Editing Challenge Baselines
Benchmark and Evaluation Protocol
Benchmark
Before November 10, 2026, participants may use the publicly available MMAE test set, comprising 2,000 examples and 17,741 atomic rubrics, to develop and evaluate their models and agent systems. The dataset is available on Hugging Face.
On November 10, 2026, the organizers will release the previously unreleased challenge test set and the submission software development kit (SDK) for uploading results. The leaderboard will open for submissions on the same date. The Single Model Track and the Agent Track will each use 500 previously unreleased test examples, and the two tracks will be ranked independently. These examples are constructed through the same MMAE pipeline and manually annotated and verified. Track 1 will be evaluated on examples from any audio modality in the MMAE single complexity category; the Agent Track will be evaluated across any audio modality and any complexity category. Test inputs and editing instructions will be provided for inference, while the rubrics remain private until the official results are finalized. The complete final set and its rubrics will be released after the competition.
Submission Format
For each test item, participants submit the edited audio file and a JSONL record that maps the sample ID to its relative audio path:
{"id": "<sample_id>", "audio_path": "audio/<sample_id>.wav"}
The audio files and JSONL manifest are packaged together and uploaded to the challenge leaderboard. The two tracks are ranked independently.
Evaluation Metrics
Each rubric is evaluated three times by a frozen Qwen3-Omni model serving as the judge, with shuffled answer choices and majority voting [17]. The leaderboard reports:
- Instruction Following Rate (IFR): the average score over rubrics that verify whether the requested edits were correctly executed.
- Consistency Rate (CR): the average score over rubrics that verify whether unrelated audio content and quality were preserved.
- Exact Match Rate (EMR): the proportion of samples for which all instruction-following and consistency rubrics are satisfied.
Systems are ranked primarily by Overall EMR, with ties broken first by Overall IFR and then by Overall CR.
Registration and Leaderboard
The registration deadline has been extended from October 1 to October 8, 2026. To participate, please complete the registration form. Register early to receive the latest challenge updates.
Before submissions open on November 10, 2026, the organizers will send a form to collect and confirm each team’s final member list, including supervisors and team leaders. Except in special circumstances, changes to team membership will not be permitted once submissions open. Each person may participate in both tracks, but may be listed on only one team per track.
Paper Submission
The top three teams in each track will be invited to submit a two-page ICASSP 2027 paper and present their work in the Grand Challenge session.
Sponsorship
The challenge prize pool is sponsored by Tencent, with a total of USD 7,000. The top three teams in the Single Model Track and the Agent Track will be awarded separately, with the following prizes in each track:
- First Prize (1st place): USD 2,000
- Second Prize (2nd place): USD 1,000
- Third Prize (3rd place): USD 500
The organizers thank Tencent for supporting this challenge and research in audio editing.
Contact
We have a Slack workspace and a WeChat group for real-time communication. For private questions, or if an invitation link or QR code has expired, please contact Zhikang Niu or Wenming Tu.
Organizers
References
- Peng, Puyuan, et al. "VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild." Proc. ACL (2024).
- Zhang, Yixiao, et al. "MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models." arXiv:2402.06178 (2024).
- Wang, Yuancheng, et al. "AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models." Proc. NeurIPS (2023).
- Tao, Ye, et al. "MMEdit: A Unified Framework for Multi-Type Audio Editing via Audio Language Model." arXiv:2512.20339 (2025).
- Tian, Zeyue, et al. "Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing." Proc. SIGGRAPH (2026).
- Lan, Zitong, et al. "Guiding Audio Editing with Audio Language Model." Proc. NeurIPS (2025).
- Qiang, Chunyu, et al. "UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions." Proc. ACL (2026).
- Gong, Junmin, et al. "ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation." arXiv:2602.00744 (2026).
- Li, Zhaoqing, et al. "UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion." arXiv:2605.31530 (2026).
- Chen, Junyang, et al. "CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models." arXiv:2601.05329 (2026).
- Zhang, Dong, et al. "MiMo-Audio: Audio Language Models are Few-Shot Learners." arXiv:2512.23808 (2025).
- Chen, William, et al. "AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing." arXiv:2602.17097 (2026).
- Wang, Hankun, et al. "dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model." arXiv:2608.02673 (2026).
- Ma, Ziyang, et al. "MMAE: A Massive Multitask Audio Editing Benchmark." arXiv:2606.07229 (2026).
- Yan, Chao, et al. "Step-Audio-EditX Technical Report." arXiv:2511.03601 (2025).
- Yan, Canxiang, et al. "Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation." arXiv:2511.05516 (2025).
- Xu, Jin, et al. "Qwen3-Omni Technical Report." arXiv:2509.17765 (2025).
- Ma, Ziyang, et al. "AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing." arXiv:2609.08936 (2026).
- Shi, Bowen, et al. "SAM Audio: Segment Anything in Audio." arXiv:2512.18099 (2025).
Follow us on GitHub for updates: @Audio-Editing-Challenge