AI Audio Separation Integration: Demucs, MDX-Net, Spleeter

AI Audio Source Separation Integration for Business

AI Development Areas

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1441
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    998
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    714
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1006

AI Audio Source Separation Integration for Business

Imagine you have a concert recording where vocals blend with guitar and drums. Making a clean karaoke track without AI is impossible. Old methods (ICA, NMF) produce artifacts, and manual processing takes hours. We solve this with modern neural networks. Our experience: 5+ years in audio processing, over 30 projects implementing source separation in media and music production. We guarantee separation quality and adherence to your deadlines.

Source separation—extracting individual sound sources from a mixed signal. It is used in music production (stems), speech processing (removing background music), video post-production, and remastering archival recordings.

What Problems We Solve

Low separation quality. Old methods (ICA, NMF) produce strong artifacts. Modern deep learning models—Demucs, Spleeter, MDX-Net—achieve SDR > 9 dB, meaning clean separation without noticeable noise.

Processing speed. For batch processing of hundreds of tracks, performance is critical. Spleeter runs 100× faster than real-time on GPU, Demucs at 1.5×. We optimize pipelines for your hardware.

Integration into existing workflows. API on FastAPI, batch processing via queues, support for popular formats—we implement it all turnkey.

How to Choose an Audio Separation Model

Choice depends on three factors: target stems, required quality, and speed.

Model Comparison Table
Model Separation Type Quality (SDR) Speed
Demucs v4 (htdemucs) Vocals/drums/bass/other 9.0 dB 1.5× realtime on GPU
Spleeter (Deezer) 2/4/5 stems 6.8 dB 100× realtime
Open-Unmix (UMX) 4 stems 7.2 dB 10× realtime
MDX-Net Competition (MDX Challenge) 9.5 dB 2× realtime
BS-RoFormer SOTA 10.1 dB 0.8× realtime

SDR (Signal-to-Distortion Ratio) is the main metric: higher = cleaner separation.

Why Demucs v4 is the Best Choice for Production

Demucs v4 (htdemucs) offers the best balance of quality and speed among open-source solutions. It is trained on large datasets and works reliably across genres. Below is a comparison of models by latency (processing time for 1 minute of audio on an A100 GPU):

Model Latency (sec) VRAM Usage
Demucs v4 0.8 2.1 GB
MDX-Net 1.2 3.8 GB
Spleeter 0.1 1.0 GB

For production, we recommend Demucs v4 or its lightweight version htdemucs_ft.

What Integration of Demucs into Your Pipeline Provides

We use Demucs v4 in production. Below is an example inference class:

import torch from demucs.pretrained import get_model from demucs.apply import apply_model from demucs.audio import AudioFile, save_audio import torchaudio class AudioSourceSeparator: def __init__(self, model_name: str = "htdemucs"): self.model = get_model(model_name) self.model.eval() if torch.cuda.is_available(): self.model.cuda() def separate( self, audio_path: str, output_dir: str, stems: list[str] = None # None = all stems ) -> dict[str, str]: """Separate track into stems, return file paths""" wav = AudioFile(audio_path).read( streams=0, samplerate=self.model.samplerate, channels=self.model.audio_channels ) ref = wav.mean(0) wav = (wav - ref.mean()) / ref.std() sources = apply_model( self.model, wav[None], device="cuda" if torch.cuda.is_available() else "cpu", progress=True, num_workers=2 )[0] sources = sources * ref.std() + ref.mean() result = {} available_stems = self.model.sources # ['drums', 'bass', 'other', 'vocals'] target_stems = stems or available_stems for stem, source in zip(available_stems, sources): if stem in target_stems: output_path = f"{output_dir}/{stem}.wav" save_audio(source, output_path, samplerate=self.model.samplerate) result[stem] = output_path return result 

Separating Speech from Background Music

For content processing, we use the htdemucs_ft model (fine-tuned on vocals). Example:

from demucs.pretrained import get_model class SpeechFromMusicExtractor: """Extract speech from video""" def __init__(self): self.model = get_model("htdemucs_ft") async def process_video_audio( self, video_path: str, output_speech: str, output_music: str ) -> dict: import subprocess audio_path = video_path.replace(".mp4", "_audio.wav") subprocess.run([ "ffmpeg", "-i", video_path, "-ac", "2", "-ar", "44100", "-vn", audio_path ], check=True) stems = self.separate(audio_path, output_dir="/tmp/stems") speech_stems = ["vocals"] music_stems = ["drums", "bass", "other"] return { "speech": stems.get("vocals"), "music_components": {k: stems[k] for k in music_stems if k in stems} } 

Our Work Process

  1. Task analysis – determine target stems, quality and speed requirements.
  2. Model selection – choose optimal architecture (Demucs, MDX-Net, Spleeter).
  3. Integration – embed the model into your pipeline (API, batch, real-time).
  4. Testing – evaluate metrics (SDR, WER) on your data.
  5. Deployment – deploy on GPU/CPU, set up monitoring.

What's Included in the Work

  • Model selection and adaptation.
  • API or batch handler implementation.
  • Operational documentation.
  • Team training (1–2 hours).
  • One month of technical support.

Typical Applications

Music production: remixing—isolate drums or bass for rework; karaoke—remove vocals, keep instrumental; mastering stems—process each layer independently.

Content and media: remove background music before STT—WER drops from 18% to 4%; remaster archival recordings—separation + denoise each stem; video localization—isolate speech, replace with dubbing.

Post-production: ADR (Automated Dialogue Replacement)—clean vocals for replacing lines; music scoring—extract music for reuse.

Limitations and Nuances

Demucs performs worse when:

  • Very loud percussion over speech (SDR drops 2–3 dB).
  • Low-quality mono recordings (< 22 kHz).
  • Complex polyphonic overlays (4+ sources simultaneously).

For maximum vocal quality, use mdx_extra or htdemucs_ft. For speed in batch mode, use Spleeter (10–15× faster than Demucs on CPU).

Timeline and Cost

Integration of Demucs into a media file processing pipeline takes 1–2 weeks. A full service with queue and web interface takes 3–4 weeks. Basic integration starts at $1,200; full service with queue and web interface starts at $3,500. Pricing is determined individually after analyzing your requirements. Contact us for a project evaluation. Get a consultation on AI separation implementation—we'll help you choose the right solution for your tasks.

Sources

  • Demucs: Hybrid Spectrogram and Waveform Source Separation, Defossez et al., 2021
  • Spleeter: A Fast and State-of-the-Art Music Source Separation Tool, Hennequin et al., 2019
  • MDX-Net: Music Demixing Challenge, 2021