AI Audio Source Separation Integration for Business
Imagine you have a concert recording where vocals blend with guitar and drums. Making a clean karaoke track without AI is impossible. Old methods (ICA, NMF) produce artifacts, and manual processing takes hours. We solve this with modern neural networks. Our experience: 5+ years in audio processing, over 30 projects implementing source separation in media and music production. We guarantee separation quality and adherence to your deadlines.
Source separation—extracting individual sound sources from a mixed signal. It is used in music production (stems), speech processing (removing background music), video post-production, and remastering archival recordings.
What Problems We Solve
Low separation quality. Old methods (ICA, NMF) produce strong artifacts. Modern deep learning models—Demucs, Spleeter, MDX-Net—achieve SDR > 9 dB, meaning clean separation without noticeable noise.
Processing speed. For batch processing of hundreds of tracks, performance is critical. Spleeter runs 100× faster than real-time on GPU, Demucs at 1.5×. We optimize pipelines for your hardware.
Integration into existing workflows. API on FastAPI, batch processing via queues, support for popular formats—we implement it all turnkey.
How to Choose an Audio Separation Model
Choice depends on three factors: target stems, required quality, and speed.
Model Comparison Table
| Model | Separation Type | Quality (SDR) | Speed |
|---|---|---|---|
| Demucs v4 (htdemucs) | Vocals/drums/bass/other | 9.0 dB | 1.5× realtime on GPU |
| Spleeter (Deezer) | 2/4/5 stems | 6.8 dB | 100× realtime |
| Open-Unmix (UMX) | 4 stems | 7.2 dB | 10× realtime |
| MDX-Net | Competition (MDX Challenge) | 9.5 dB | 2× realtime |
| BS-RoFormer | SOTA | 10.1 dB | 0.8× realtime |
SDR (Signal-to-Distortion Ratio) is the main metric: higher = cleaner separation.
Why Demucs v4 is the Best Choice for Production
Demucs v4 (htdemucs) offers the best balance of quality and speed among open-source solutions. It is trained on large datasets and works reliably across genres. Below is a comparison of models by latency (processing time for 1 minute of audio on an A100 GPU):
| Model | Latency (sec) | VRAM Usage |
|---|---|---|
| Demucs v4 | 0.8 | 2.1 GB |
| MDX-Net | 1.2 | 3.8 GB |
| Spleeter | 0.1 | 1.0 GB |
For production, we recommend Demucs v4 or its lightweight version htdemucs_ft.
What Integration of Demucs into Your Pipeline Provides
We use Demucs v4 in production. Below is an example inference class:
import torch
from demucs.pretrained import get_model
from demucs.apply import apply_model
from demucs.audio import AudioFile, save_audio
import torchaudio
class AudioSourceSeparator:
def __init__(self, model_name: str = "htdemucs"):
self.model = get_model(model_name)
self.model.eval()
if torch.cuda.is_available():
self.model.cuda()
def separate(
self,
audio_path: str,
output_dir: str,
stems: list[str] = None # None = all stems
) -> dict[str, str]:
"""Separate track into stems, return file paths"""
wav = AudioFile(audio_path).read(
streams=0,
samplerate=self.model.samplerate,
channels=self.model.audio_channels
)
ref = wav.mean(0)
wav = (wav - ref.mean()) / ref.std()
sources = apply_model(
self.model,
wav[None],
device="cuda" if torch.cuda.is_available() else "cpu",
progress=True,
num_workers=2
)[0]
sources = sources * ref.std() + ref.mean()
result = {}
available_stems = self.model.sources # ['drums', 'bass', 'other', 'vocals']
target_stems = stems or available_stems
for stem, source in zip(available_stems, sources):
if stem in target_stems:
output_path = f"{output_dir}/{stem}.wav"
save_audio(source, output_path, samplerate=self.model.samplerate)
result[stem] = output_path
return result
Separating Speech from Background Music
For content processing, we use the htdemucs_ft model (fine-tuned on vocals). Example:
from demucs.pretrained import get_model
class SpeechFromMusicExtractor:
"""Extract speech from video"""
def __init__(self):
self.model = get_model("htdemucs_ft")
async def process_video_audio(
self,
video_path: str,
output_speech: str,
output_music: str
) -> dict:
import subprocess
audio_path = video_path.replace(".mp4", "_audio.wav")
subprocess.run([
"ffmpeg", "-i", video_path,
"-ac", "2", "-ar", "44100",
"-vn", audio_path
], check=True)
stems = self.separate(audio_path, output_dir="/tmp/stems")
speech_stems = ["vocals"]
music_stems = ["drums", "bass", "other"]
return {
"speech": stems.get("vocals"),
"music_components": {k: stems[k] for k in music_stems if k in stems}
}
Our Work Process
- Task analysis – determine target stems, quality and speed requirements.
- Model selection – choose optimal architecture (Demucs, MDX-Net, Spleeter).
- Integration – embed the model into your pipeline (API, batch, real-time).
- Testing – evaluate metrics (SDR, WER) on your data.
- Deployment – deploy on GPU/CPU, set up monitoring.
What's Included in the Work
- Model selection and adaptation.
- API or batch handler implementation.
- Operational documentation.
- Team training (1–2 hours).
- One month of technical support.
Typical Applications
Music production: remixing—isolate drums or bass for rework; karaoke—remove vocals, keep instrumental; mastering stems—process each layer independently.
Content and media: remove background music before STT—WER drops from 18% to 4%; remaster archival recordings—separation + denoise each stem; video localization—isolate speech, replace with dubbing.
Post-production: ADR (Automated Dialogue Replacement)—clean vocals for replacing lines; music scoring—extract music for reuse.
Limitations and Nuances
Demucs performs worse when:
- Very loud percussion over speech (SDR drops 2–3 dB).
- Low-quality mono recordings (< 22 kHz).
- Complex polyphonic overlays (4+ sources simultaneously).
For maximum vocal quality, use mdx_extra or htdemucs_ft. For speed in batch mode, use Spleeter (10–15× faster than Demucs on CPU).
Timeline and Cost
Integration of Demucs into a media file processing pipeline takes 1–2 weeks. A full service with queue and web interface takes 3–4 weeks. Basic integration starts at $1,200; full service with queue and web interface starts at $3,500. Pricing is determined individually after analyzing your requirements. Contact us for a project evaluation. Get a consultation on AI separation implementation—we'll help you choose the right solution for your tasks.
Sources
- Demucs: Hybrid Spectrogram and Waveform Source Separation, Defossez et al., 2021
- Spleeter: A Fast and State-of-the-Art Music Source Separation Tool, Hennequin et al., 2019
- MDX-Net: Music Demixing Challenge, 2021







