Processed on this device

AI Vocal Remover: Remove the Voice from a Video and Keep the Music

Drop in a video or a song and the AI separates the sung or spoken voice from the soundtrack behind it. The music, the score and the sound effects stay exactly where they were, and you get the video back with the picture untouched. On a computer it runs free in your browser, on this device: nothing is uploaded and no account is needed. For longer files, or a browser that cannot run the model, switch to Cloud: your first 10 credits are free, which is 10 minutes here.

Remove Audio mascot, friendly headphone character. AI audio tool that runs free in your browser where it can, or on our servers for longer files.

Drop a video or audio file or browse your device

MP4, MOV, MKV, WebM, AVI, MP3, WAV, M4A and more

Up to 20 files at once

Privacy-first No uploads Processed locally

This browser cannot run the AI model on its own: it needs WebGPU and at least 8 GB of memory. Switch to Cloud to run this file.

Free on this device where the model can run. In the Cloud, you pick how long the result is kept.

Bigger files? Run it in the Cloud

This tool runs on this device for free. For a file over the device limit or a long recording, the same tool runs on our servers, with up to 5 GB per file and 6 hours of audio, faster, and with your files kept for up to 7 days, or as long as you choose with any pack or plan. From 1 credit a minute, with 10 credits free when you create an account.

Quick answer

How to remove the voice from a video and keep the music

Drop in the video and the AI separates the sung or spoken voice from everything else. On this device that runs free in your browser with no upload. You get the same video back with the soundtrack intact and the voice gone. No editor, and no separate audio file to re-import.

  1. Drop in an MP4, MOV, MKV or WebM, or an MP3 or WAV.
  2. The AI splits the voice away from the music and effects.
  3. Download the video with the music still playing.

Why use this instead of a video editor

Editors can lower a track or cut a range, but they cannot pull one voice out of a soundtrack that is already mixed. This separates the parts, so you keep everything you wanted and lose only the voice.

Video in, video out

Most vocal removers hand you an audio stem and leave you to rebuild the video yourself. Drop a video here and you get a video back, with the picture untouched.

Works on a finished mix

You do not need the original stems or a project file. Point it at a clip where the voice and music are already mixed together and it separates them.

Keeps the sound effects

It targets the voice, not everything that is not music. Footsteps, engines, crowd noise and the score all stay where they were.

No editing skill needed

There is nothing to line up on a timeline. Drop the file in, wait, download. The whole job is a few clicks and takes minutes, not an evening.

Handles long recordings

A full gameplay session or a lecture recording works the same way a 30 second clip does: up to 15 minutes runs in your browser, and anything longer can run in the Cloud with an account. You are not capped at a short preview.

Free in your browser

This device runs the AI model on your own computer, so it costs nothing: no account, no upload, no credits, as many files as you like. The Cloud, for long files, starts every new account with 10 credits for free, which is 10 minutes of this tool. No card needed.

How it works, step by step

Four steps from a clip with talking over it to a clip with just the music.

  1. 01

    Add your file

    Drag a video or audio file onto the box above, or click to browse. MP4, MOV, MKV and WebM all work, as do MP3, WAV and M4A.

  2. 02

    Pick This device or Cloud

    This device runs the AI in your browser: free, no account, nothing uploaded, on a browser with WebGPU and at least 8 GB of memory, up to 15 minutes of audio per file. The Cloud runs it on our servers for longer files: sign in, and your first 10 credits are free.

  3. 03

    Let the model separate the audio

    It listens to the whole track and pulls the sung and spoken voice away from the music, the effects and the room sound behind it.

  4. 04

    Download the result

    When it is done, download the video with the music kept and the voice removed. On this device your file has not left it; in the Cloud the files are deleted on the schedule you picked.

What people use it for

This tool pulls the vocals and spoken voices out of a track and leaves the music and soundtrack behind. Here are the jobs people bring to it.

Drop the commentary

Take a gameplay clip or a reaction video and strip the talking while the background track keeps playing. You get the music bed without the voice over the top, ready to re-edit for a Reel or Short at 9:16.

Make a karaoke version

Feed in a song and pull the lead vocal out to leave the instrumental. You get a backing track you can sing over, with the drums, bass, and melody intact.

Remove narration from a clip

You have a tutorial or a travel video where the voice walks over the music underneath. Drop out the narration and keep the soundtrack, so the clip still has its mood without anyone speaking.

Clean bed for a remix

Strip the vocals so you are left with a clean instrumental to build on. Export it and pull the music into your DAW or your edit without a voice cutting through the mix.

Keep the film score

On a movie or trailer clip, drop the dialogue and hold on to the score. You keep the strings and the swell while the spoken lines come out, which is handy when you only want the music for a project.

Instrumental for video

Turn a track with vocals into a clean instrumental to sit under a 1:1 Instagram post or a talking-head clip. The music plays underneath without competing words, so your own voice or captions stay clear.

What it handles well, and what it does not

Works well onStruggles with
Commentary or narration over a music bedWhispered or very quiet speech buried under a loud mix
A clear lead vocal over an instrumentalHeavily processed or vocoded vocals that sound like an instrument
Dialogue over a film scoreRecordings where the voice and music share the same frequencies
Reaction and gameplay clipsVery low bitrate audio that was already badly compressed
How it compares

Against the other ways to do this

ApproachWhat you actually get
Muting the clipSilence. The music goes too, which is usually not what you wanted.
A karaoke or vocal remover siteUsually an MP3 stem. You still have to rebuild the video around it in an editor.
The phase inversion trick in AudacityWorks only on some stereo mixes, and it hollows out the music. Most modern mixes defeat it.
Premiere, DaVinci or CapCutYou can duck or cut the whole track, but nothing in the timeline separates one voice from a finished mix.
This toolThe same video back, picture untouched, voice gone and the soundtrack still playing.

The honest summary: if your voice and music are still on separate tracks in a project you own, an editor is the better answer. This is for finished files where they are already mixed together.

Wanted the opposite? The background music remover keeps your talking and drops the music, and the AI voice isolator does the same job on any recording.

If something does not sound right

Frequently asked questions

Yes, in your browser. Pick This device and the AI model runs on your own computer: no account, no upload, no credits, and it stays free however many files you run. It needs a browser with WebGPU and at least 8 GB of memory, handles up to 15 minutes of audio per file, and downloads the model once, on the first run. The Cloud is for longer files and for devices that cannot run the model: every new account starts with 10 credits for free, which is 10 minutes of this tool, no card needed. The Cloud costs 1 credit per minute, and every started minute counts, so a 2:30 clip is billed as 3 minutes.

Last reviewed September 2026.