txscribe

Remove background noise

Drop in a recording and get it back with the hum, hiss, traffic and chatter turned down. The cleaning runs in your browser: nothing is uploaded unless you press Transcribe afterwards, and there’s no account.

How it works

The noise is removed by DeepFilterNet 3, an open-source neural network made for speech. It listens to the recording ten milliseconds at a time and turns down, band by band, what it judges not to be voice.

LICENSE-DeepFilterNet-MIT.txt · LICENSE-DeepFilterNet-APACHE.txt · THIRD-PARTY-NOTICES.txt

Your browser decodes the file, the model runs on your device’s processor, and the cleaned audio is saved from this page. The first time you choose a recording, the page downloads the model and the code that runs it from txscribe.ai — about 20 MB once unpacked — and your browser keeps it for next time.

The cleaned file lines up with the original to the sample, so you can switch between them in the middle of a word and hear exactly what changed.

What it handles, and what it doesn’t

Hum, hiss, air conditioning, rain, keyboards, traffic
Turned down the most. In our tests, in the pauses between words, each of these fell to a small fraction of its level. Under the words, where the noise was as loud as the voice, a little of the voice went with it.
Other people’s voices in the background
It depends on the voices. In our tests a restaurant’s murmur fell as far as traffic did; a classroom, where single voices stand out, fell less than half as far as traffic. The model is built to keep voices.
Another person talking near the microphone
Made quieter, not removed: in our tests a second voice came out at about half its level, and the main voice lost a little too. The model cannot tell which speaker you wanted.
A voice quieter than the noise over it
The hardest case: it can go with the noise. In our test a distant voice under a louder room came out at about a tenth of its level. Listen to those parts before you use the result.
A recording that is already clean
Comes back almost unchanged.

Limits and formats

Up to 10 minutes and 500 MB

What it opens
MP3, M4A, AAC, WAV, FLAC, OGG, Opus and WebM audio, and the sound of MP4 and MOV videos: each of these decoded in every browser engine we tested. Anything else depends on your browser; if it can’t decode a file, the page says so.
What you get
A WAV file: 16-bit, 48 kHz, one channel — about 5.5 MB a minute. A stereo recording is mixed to mono first. From a video you get its sound, not a new video.

Privacy

The recording stays on your device: it is decoded, cleaned and saved here, and none of it is sent anywhere. We count which steps of the page are used — a file chosen, cleaned, downloaded — with no file name, length or sound attached. Only Transcribe uploads the original, and it says so beside the button.

Questions

Is it free?

Yes, with no account, no watermark and nothing to install. The limits are the file’s: up to 10 minutes and 500 MB.

Is my recording uploaded?

Cleaned in this browser. Nothing is uploaded unless you press Transcribe afterwards.

Why do some words sound thinner afterwards?

Where a voice and the noise behind it are about as loud as each other, the model has to decide, moment by moment, what is voice, and some of the voice is turned down with the noise. A word quieter than the noise over it can all but disappear. Switch between Before and After on the parts that matter before you use the result.

Can it clean the sound of a video?

Yes: choose the video and you get its sound back, cleaned, as a WAV — not a new video.

Why does Transcribe use the original recording, not the cleaned one?

The cleaned copy is made for listening. Cleaning changes the voice a little as well as the noise, and we do not claim it makes a transcript more accurate — so Transcribe sends the recording as you chose it.

Which browsers does it work in?

Recent ones: Chrome or Edge, Firefox 114 or later, or Safari 16.4 or later. Older browsers lack what it runs on (WebAssembly with SIMD, and workers that load modules), and the page says so if yours can’t run it. It runs on your device’s processor and in its memory, so a slower device takes longer, and a phone can run out of memory on a long recording; the page shows the time left.

Need the words, not just the sound?

txscribe transcribes recordings in 82 languages, with a timestamp on every word, so any line plays back the moment it came from. The free trial takes one recording, up to minute 10, with no account.

Start free