AI Text-to-Speech (Inflect Nano v2)

Convert text to English speech with the tiny Inflect Nano v2 model. Runs privately in your browser with a download of about 16 MB.

This model requires downloading ~16.21 MB of data on first use. All processing runs locally in your browser. View on Hugging Face


Other Audio Tools


Inflect v2 AI Text-to-Speech in Your Browser

Inflect v2 converts English text into a 24 kHz spoken waveform entirely in your browser. Choose Inflect Nano v2 for the smallest download or Inflect Micro v2 for the stronger quality-focused model. Both versions run locally after their model files have loaded, so your text is not sent to a speech-generation server.

What is the difference between Inflect Micro v2 and Inflect Nano v2?

Inflect Nano v2 is the portability-focused option with approximately 3.96 million parameters and a 16.21 MB model download. Inflect Micro v2 uses approximately 9.36 million parameters and 37.75 MB of model data to prioritize speech quality. Both use the same English voice, output format, and speed control.

Does Inflect text-to-speech run locally?

Yes. The model executes through ONNX Runtime Web, while an eSpeak-ng phonemizer prepares the English text. Downloads come directly to your browser, and synthesis runs on your device without an account or API key. The first generation takes longer because the browser must download and initialize the model.

Can I change the voice or language?

Inflect v2 provides one fixed synthetic male English voice. It does not support voice cloning, reference audio, multiple speakers, or additional languages. For a choice of voices and multilingual synthesis, use Supertonic TTS 2.

How does the speech speed control work?

Use the speed slider to select a value from 0.5× to 2×. Lower values produce slower speech, while higher values shorten the predicted speech duration. The model generates the waveform at the selected pace instead of changing the playback speed afterward.

How are numbers, dates, and abbreviations pronounced?

The browser frontend expands common English formats before phonemization. It handles text such as dates, times, money amounts, ordinals, acronyms, and frequently used abbreviations so they are read as words rather than isolated symbols. Uncommon names, homographs, or specialist notation can still require spelling adjustments.

Can I convert long text to speech?

Long input is split at sentence and punctuation boundaries, synthesized in manageable chunks, and joined with short pauses. This supports paragraphs that exceed a single model pass while keeping sentence transitions clear. Very long documents may still take time because every chunk is generated locally.

How can I save the generated speech?

After generation finishes, listen in the built-in audio player or download the result as a WAV file. You can turn that audio into a shareable visual clip with the Audio to Video Visualizer, or check its spoken content with the AI Audio Transcriber.

Navigator

Quickly navigate to any tool