Text to Speech
Turn text into natural speech with 16 American and British AI voices. The model runs in your browser, your text stays private, and use is unlimited.
The voice model runs inside your browser, so your text stays on your device and there are no word limits. The first generation downloads the model into your browser cache (about 90 MB, one time). After that it works even offline.
record_voice_overVoice
lightbulbTip
Heart and Bella are the highest rated American voices, and Emma leads the British set. Long texts are read sentence by sentence, and the progress counter shows the audio growing while you wait.
What is the Text to Speech?
A text to speech tool reads written text aloud with a synthetic voice. This one runs Kokoro, an 82 million parameter speech model with Apache licensed weights, directly in your browser through WebGPU or WebAssembly. Because your own device does the work, there are no word limits, no accounts, and no usage fees, and the text you paste never reaches a server. Long passages are read sentence by sentence and stitched into a single audio file you can play on the page or download as WAV.
How to use the Text to Speech
- 1
Paste your text
Type or paste up to 20,000 characters. Articles, scripts, study notes, and announcements all work.
- 2
Pick a voice and speed
Choose between American and British voices, female and male, each labeled with its quality grade from the model card. Adjust speed from 0.5x to 2x.
- 3
Generate
The first run downloads the voice model into your browser cache, about 90 MB one time. Progress shows the audio growing sentence by sentence.
- 4
Listen and download
Play the result right on the page, then download it as a WAV file that any editor or player accepts.
Frequently Asked Questions
Is it really free and unlimited?
Yes. The speech model runs on your own device, so there is nothing for us to meter. Generate an hour of audio if you like; the only limit is your hardware's speed.
Is my text uploaded anywhere?
No. After the one-time model download, generation happens entirely in your browser. You can watch the Network tab while it runs, and it even works offline once the model is cached.
Which languages and voices are available?
The current Kokoro model speaks English with 16 curated American and British voices, female and male. Heart and Bella carry the highest quality grades on the model card.
Can I use the audio commercially?
The Kokoro model is released under the Apache 2.0 license, which permits commercial use, and we claim no rights over audio you generate. For high-stakes projects, review the model license yourself.
Why is the first generation slow?
The first run downloads the 90 MB model and warms it up. After that, devices with WebGPU generate speech many times faster than real time, and the model loads from cache instantly.
Can I download MP3 instead of WAV?
The tool exports WAV, which every player, editor, and video tool accepts without quality loss. If you need MP3 for size reasons, any audio converter can compress the WAV afterwards.