How to use
- Paste text.
- Pick a voice and speed, then download the model (first time only) and generate. Browser voices play instantly but cannot be downloaded.
- Listen, then download WAV or MP3.
Worked example
The 40-word sample text with the Heart voice at 1.0× becomes 16.5 seconds of speech, a 771 KB WAV.
Supported formats and limits
| Input | Text up to 5,000 characters |
|---|---|
| Output | WAV (24 kHz mono), MP3, Browser voices: playback only |
| Limits | First use downloads about 93 MB of model files from Hugging Face (cached afterwards). English voices (US and UK). |
| Engine | Kokoro-82M (Apache-2.0, 8-bit ONNX) via kokoro-js in a Web Worker, sentence by sentence; or the browser's speechSynthesis voices |
Limitations
- Unusual names and abbreviations may be mispronounced; spell them phonetically.
- Only English is supported by the bundled voices.
Questions
Can I download the audio?
Yes with Kokoro: the speech is generated as a 24 kHz mono WAV, or MP3. Browser voices play instantly but are playback only, because browsers do not hand that audio to the page.
What is downloaded the first time?
About 93 MB of Kokoro model files from Hugging Face, plus a 0.5 MB style file per voice, cached for later visits. Your text is not sent.
Privacy
On-device model. Processing runs in this browser. The open-source model files are downloaded once from the model host (Hugging Face) and cached; your content is not uploaded.
- Hugging Face: Downloads of the open-source model and voice files when you ask for them. No text or audio is sent.
See the privacy policy for how toolsdocks handles data.