Skip to content

Stem separation

Three separation models, one task. Miso offers Split into stems and runs whichever of the three you have installed.

Three entries in the tool list would be three ways to do one thing, so there is one.

Model Output VRAM
Mel-Band RoFormer vocals, instrumental 4 to 6 GB
BS-RoFormer vocals, instrumental 4 to 6 GB
HTDemucs vocals, drums, bass, other 3 to 4 GB

The two RoFormer models give the cleanest vocal extraction with the least bleed. HTDemucs is the one to install when you want drums and bass on their own.

This task has no fields at all, and that is not an omission. All three vendored specs carry an empty request option list, and every probe run sent nothing but the audio.

You pick the model and the source. That is the whole form.

All three refuse anything that is not 44.1 kHz, before they start any work.

Miso handles this for generated takes. Every take audio.cpp writes is 48 kHz, so the worker converts before sending. The task declares the rate it needs rather than the worker knowing which families are fussy.

An imported file at some other rate is the case you have to fix yourself, in Audio tools.

A converted source skips the staged upload cache in both directions, because that cache is keyed by asset and backend and cannot tell a converted copy from the original.

Stems, not takes. They are filed together under the job that made them and appear on the Stems screen, where you can play them, mix them back, download the set as a zip, or convert a vocal stem to another voice.

See Splitting into stems.