AI Vocal Remover: Remove the Voice from a Video and Keep the Music
Drop in a video or a song and the AI separates the sung or spoken voice from the soundtrack behind it. The music, the score and the sound effects stay exactly where they were, and you get the video back with the picture untouched. On a computer it runs free in your browser, on this device: nothing is uploaded and no account is needed. For longer files, or a browser that cannot run the model, switch to Cloud: your first 10 credits are free, which is 10 minutes here.
Drop a video or audio file or browse your device
MP4, MOV, MKV, WebM, AVI, MP3, WAV, M4A and more
Up to 20 files at once
This browser cannot run the AI model on its own: it needs WebGPU and at least 8 GB of memory. Switch to Cloud to run this file.
Free on this device where the model can run. In the Cloud, you pick how long the result is kept.
Bigger files? Run it in the Cloud
This tool runs on this device for free. For a file over the device limit or a long recording, the same tool runs on our servers, with up to 5 GB per file and 6 hours of audio, faster, and with your files kept for up to 7 days, or as long as you choose with any pack or plan. From 1 credit a minute, with 10 credits free when you create an account.
How to remove the voice from a video and keep the music
Drop in the video and the AI separates the sung or spoken voice from everything else. On this device that runs free in your browser with no upload. You get the same video back with the soundtrack intact and the voice gone. No editor, and no separate audio file to re-import.
- Drop in an MP4, MOV, MKV or WebM, or an MP3 or WAV.
- The AI splits the voice away from the music and effects.
- Download the video with the music still playing.
Why use this instead of a video editor
Editors can lower a track or cut a range, but they cannot pull one voice out of a soundtrack that is already mixed. This separates the parts, so you keep everything you wanted and lose only the voice.
Video in, video out
Most vocal removers hand you an audio stem and leave you to rebuild the video yourself. Drop a video here and you get a video back, with the picture untouched.
Works on a finished mix
You do not need the original stems or a project file. Point it at a clip where the voice and music are already mixed together and it separates them.
Keeps the sound effects
It targets the voice, not everything that is not music. Footsteps, engines, crowd noise and the score all stay where they were.
No editing skill needed
There is nothing to line up on a timeline. Drop the file in, wait, download. The whole job is a few clicks and takes minutes, not an evening.
Handles long recordings
A full gameplay session or a lecture recording works the same way a 30 second clip does: up to 15 minutes runs in your browser, and anything longer can run in the Cloud with an account. You are not capped at a short preview.
Free in your browser
This device runs the AI model on your own computer, so it costs nothing: no account, no upload, no credits, as many files as you like. The Cloud, for long files, starts every new account with 10 credits for free, which is 10 minutes of this tool. No card needed.
How it works, step by step
Four steps from a clip with talking over it to a clip with just the music.
Add your file
Drag a video or audio file onto the box above, or click to browse. MP4, MOV, MKV and WebM all work, as do MP3, WAV and M4A.
Pick This device or Cloud
This device runs the AI in your browser: free, no account, nothing uploaded, on a browser with WebGPU and at least 8 GB of memory, up to 15 minutes of audio per file. The Cloud runs it on our servers for longer files: sign in, and your first 10 credits are free.
Let the model separate the audio
It listens to the whole track and pulls the sung and spoken voice away from the music, the effects and the room sound behind it.
Download the result
When it is done, download the video with the music kept and the voice removed. On this device your file has not left it; in the Cloud the files are deleted on the schedule you picked.
What people use it for
This tool pulls the vocals and spoken voices out of a track and leaves the music and soundtrack behind. Here are the jobs people bring to it.
Drop the commentary
Take a gameplay clip or a reaction video and strip the talking while the background track keeps playing. You get the music bed without the voice over the top, ready to re-edit for a Reel or Short at 9:16.
Make a karaoke version
Feed in a song and pull the lead vocal out to leave the instrumental. You get a backing track you can sing over, with the drums, bass, and melody intact.
Remove narration from a clip
You have a tutorial or a travel video where the voice walks over the music underneath. Drop out the narration and keep the soundtrack, so the clip still has its mood without anyone speaking.
Clean bed for a remix
Strip the vocals so you are left with a clean instrumental to build on. Export it and pull the music into your DAW or your edit without a voice cutting through the mix.
Keep the film score
On a movie or trailer clip, drop the dialogue and hold on to the score. You keep the strings and the swell while the spoken lines come out, which is handy when you only want the music for a project.
Instrumental for video
Turn a track with vocals into a clean instrumental to sit under a 1:1 Instagram post or a talking-head clip. The music plays underneath without competing words, so your own voice or captions stay clear.
What it handles well, and what it does not
| Works well on | Struggles with |
|---|---|
| Commentary or narration over a music bed | Whispered or very quiet speech buried under a loud mix |
| A clear lead vocal over an instrumental | Heavily processed or vocoded vocals that sound like an instrument |
| Dialogue over a film score | Recordings where the voice and music share the same frequencies |
| Reaction and gameplay clips | Very low bitrate audio that was already badly compressed |
Against the other ways to do this
| Approach | What you actually get |
|---|---|
| Muting the clip | Silence. The music goes too, which is usually not what you wanted. |
| A karaoke or vocal remover site | Usually an MP3 stem. You still have to rebuild the video around it in an editor. |
| The phase inversion trick in Audacity | Works only on some stereo mixes, and it hollows out the music. Most modern mixes defeat it. |
| Premiere, DaVinci or CapCut | You can duck or cut the whole track, but nothing in the timeline separates one voice from a finished mix. |
| This tool | The same video back, picture untouched, voice gone and the soundtrack still playing. |
The honest summary: if your voice and music are still on separate tracks in a project you own, an editor is the better answer. This is for finished files where they are already mixed together.
Wanted the opposite? The background music remover keeps your talking and drops the music, and the AI voice isolator does the same job on any recording.
If something does not sound right
Frequently asked questions
Yes, in your browser. Pick This device and the AI model runs on your own computer: no account, no upload, no credits, and it stays free however many files you run. It needs a browser with WebGPU and at least 8 GB of memory, handles up to 15 minutes of audio per file, and downloads the model once, on the first run. The Cloud is for longer files and for devices that cannot run the model: every new account starts with 10 credits for free, which is 10 minutes of this tool, no card needed. The Cloud costs 1 credit per minute, and every started minute counts, so a 2:30 clip is billed as 3 minutes.
Last reviewed September 2026.
More tools
Mute Video
Strip every audio track from a video. Free, in your browser, no upload.
Open toolVideo to MP3
Pull the audio out of any video and save it as an MP3. Nothing leaves your browser.
Open toolCompress Video
Shrink a video to fit Discord, WhatsApp or email size limits. Free on this device, in your browser.
Open tool