Home/ The Signal/ Tutorials/ Vocal Sample Cleaning, Tested: Ultimate Vocal Remover vs LALAL.AI vs Moises on Three Filthy Sources
Stem Separation

Vocal Sample Cleaning, Tested: Ultimate Vocal Remover vs LALAL.AI vs Moises on Three Filthy Sources

Four bars of a 1974 soul 45, ripped from a YouTube upload of somebody's turntable. Pitched down and time-stretched to 78 BPM, the phrasing sat perfectly across the bar.

A photorealistic close-up photograph of a cluttered home studio desk at night, shot at…

Four bars of a 1974 soul 45, ripped from a YouTube upload of somebody's turntable. Pitched down and time-stretched to 78 BPM, the phrasing sat perfectly across the bar. Played against a clean 48kHz drum bus, it sounded like a phone call. Surface noise, a small room, and a Rhodes chord that would not leave the vocal alone. That gap is the whole problem with vocal sample cleaning: the take you want almost never arrives in a state your track will accept, and until recently the fix meant either a restoration suite priced like a used car or an afternoon of surgical EQ that removed the life along with the hiss.

What changed is that source separation models got good enough to run on a laptop, and three different business models grew up around them. I put the same three dirty sources through all three, at the same settings depth a working producer would actually bother with, and paid attention to how each one fails rather than how it markets.

The three on the bench

Ultimate Vocal Remover (UVR) is a free, open-source desktop app that acts as a front end for downloadable model checkpoints — the MDX-Net, Demucs and VR architectures that most of this field is built on. Nothing uploads. You pick a model, pick a chunk size, and wait. On a decent GPU a four-minute file takes a couple of minutes. On a laptop CPU, go make coffee and answer some email.

LALAL.AI is the commercial browser-and-desktop option, sold in minute packs rather than a flat subscription. Alongside stem splitting it ships processing that the open-source models do not do on their own: noise cancelling with intensity levels, a de-echo pass, and a general voice cleaner. Its separation algorithm gets a new codename every so often — Lynx, as of writing — which is worth reading as "the current one" rather than as a permanent product.

Moises is subscription-shaped and built for musicians practising along to records: separation plus key and tempo detection, pitch shift, chord charts, a metronome. Sampling is a thing you can do with it rather than the thing it was designed for, and that shows up in the output.

How I decided

Six criteria, because "quality" on its own tells you nothing about whether a tool survives a real session:

  • Failure character. Every separator fails. What matters is whether it fails as a dull, dark vocal (usable, you can add air back) or as chirping high-band artefacts around the sibilance (unusable, that sizzle is permanent).
  • What it removes besides instruments. Bleed is one problem. Room reverb, tape hiss and crowd noise are three others, and they are the ones that actually ruin a vinyl or live rip.
  • Export truth. Sample rate, bit depth, mono or stereo, WAV or MP3. Check this in the export panel before you buy anything, because a beautiful separation delivered as a 44.1kHz MP3 into a 48kHz project is a resample and a generation loss you did not ask for.
  • Throughput. Fifty candidate rips is a normal crate-digging session. A tool that makes each one a five-click ceremony will not get used.
  • Cost shape over twelve months. Per-minute, per-month, or free-but-your-electricity. These are different risks, not different prices.
  • Friction. Setup time, upload time, and the small tax of deciding whether this particular sample is worth spending metered minutes on.

Test one: the vinyl 45

The soul rip. Continuous surface noise, a mid-forward mastering job, and a Rhodes doubling the vocal melody a third below — which is the hardest possible case, because the model has to separate two sources that share partials.

UVR with an MDX-Net vocal model got the cleanest raw split. The Rhodes came out almost entirely; what stayed was a faint harmonic ghost on sustained notes, which sits under a drum bus without complaint. What UVR did not do was touch the surface noise. That crackle rode along with the vocal, because as far as the model is concerned broadband noise correlates with the voice.

LALAL.AI's split left a touch more keys bleed on the long notes, but its noise cancelling is where it earned its keep. On the middle intensity the crackle dropped to a whisper and the vocal kept its top end. On the highest setting the whole thing went underwater — consonants lost their edge and the sample took on that dead, gated quality that screams "processed" the moment it's in a mix. Middle setting, every time. That is the single most useful dial in this entire category and it is the one people overcook.

Moises separated the vocal fine and left the noise, same as UVR, with slightly softer transients on the consonants. For a practice tool that is a reasonable trade. For a sample you are about to chop, softer transients mean your slice points get mushy.

Worth saying: on a lo-fi or boom-bap record, some of that crackle is the sound you came for. Clean the vocal, then put your own noise back under it where you control the level.

Test two: the live bootleg

A fan-shot festival performance from YouTube, audio effectively mono, a crowd singing along, and PA slap off the back of the field. This is the source that separates the tools.

All three pulled a vocal. None of them pulled a clean vocal, because the crowd is also vocals — the models correctly identified two thousand people as the thing to keep. UVR gave me the lead plus an eerie choir, which honestly worked for the track and I kept it. That is a lucky accident, not a feature.

LALAL.AI's de-echo helped with the PA slap, tightening the tail on the ends of phrases from a smeared quarter-note to something closer to a plate. It did not remove the reverb; it shortened it, and a residue stayed on the loudest words. That is the honest ceiling on dereverb right now across every tool I have used — it changes the size of the room, it does not delete the room.

Moises struggled most here, with the crowd bleeding through as a wash and a wobbly artefact on the held notes that only got worse when I pitched the sample down.

Test three: the phone a cappella

A singer friend's voice memo, recorded in a tiled bathroom, mildly clipped on the loud phrases. No instruments to separate at all — a pure voice isolation AI problem rather than a stem problem.

This is where the free option runs out of road. UVR has nothing to say about reverb; separation models expect a mixture, and there isn't one. LALAL.AI's de-echo and voice cleaner made a real difference, moving the take from "unusable" to "usable if I commit to it being a texture rather than a lead." The clipping stayed clipped, because nothing here reconstructs destroyed peaks. Moises landed in between, cleaner than raw but not tight.

The table

Criterion UVR LALAL.AI Moises
Raw separation on dense music Best in test Very close Softer transients
Noise removal None built in Strong, three intensities None built in
De-echo / dereverb None Shortens, does not delete None
Export Whatever the model outputs, locally, no ceiling Check the format panel per plan Format ceilings vary by tier
Throughput Batch folders, no meter anxiety Metered minutes, watch the clock App-shaped, slower per file
Cost shape Free, costs you time and hardware Pay per minute in packs Recurring, whether or not you sample
Honest negative Setup and model roulette; zero support Overcooked settings smear sibilance Tuned for practice, not chopping

Cost over twelve months

The money question is not which is cheaper, it is which failure mode you can live with. Per-minute credits penalise experimentation — you start rationing, and rationing is how you stop auditioning marginal samples, which is where the good ones hide. Subscriptions keep charging through the three months you spend mixing rather than digging. Free costs you an evening of setup, a machine that can take it, and the acceptance that when a model produces garbage there is nobody to email.

My actual rig after this: UVR does the volume work on everything, and metered minutes get spent only on the handful of rips that survive the first pass and need noise or echo work. Free for breadth, paid for depth.

The part no separator fixes

Isolating a vocal does not give you the right to use it. A separator is a technical operation on a recording that somebody owns; the master rights and the publishing rights both survive the process intact, and a clean acapella is arguably easier to identify than a buried one. Check each tool's terms for what it claims over your uploads and outputs, because that varies and it changes. Cleaning up vocal rips is a production step. Clearance is a separate job, and it is the one that ends projects.

Who this is for

If you are chopping dozens of candidates a week and you own a machine with a GPU, UVR is the correct default and the price is unbeatable. If your sources are genuinely damaged — vinyl noise, room, crowd — and your time is worth more than the credits, LALAL.AI's cleanup stages do work that the free models do not attempt. Skip Moises for sampling; it is a fine practice tool doing a job it wasn't built for, and you will feel it at the slice points.

What this didn't answer: none of it tells you how these hold up on non-English vocals, on heavily processed sources where the vocal is already drenched in effects, or on anything where the singer is doubled and panned. It also says nothing about batch scripting Demucs directly from the command line, which is where this goes once fifty files a week becomes two hundred. Start there, and start with a source you have already cleared.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
I

Imogen Hale

Music-Tech & Licensing Reporter

Imogen Hale reports on the business side of AI music — licensing terms, royalties, and copyright — reading the fine print so working creators don't get burned. More by Imogen Hale →