Ito TTS: Streaming Speech for a $5 ESP32-S3

Introduction
Lokutor (Daniel Varela) published Ito—a tiny streaming neural TTS stack aimed at a $5-class ESP32-S3 with no NPU. Primary sources: the Oct 2, 2026 blog Ito: streaming neural speech on a $5 chip, the gated HF card lokutor-ai/ito (blog also mentions ito-tts-v3), and GitHub lokutor-ai/ito.
Headline numbers: ~4.4M parameters / ~4.9 MB int8, two voices (female D / male G), 24 kHz, streaming with a claimed ~25 ms first audio chunk then ~100 ms chunks. English only. You can run it on a laptop today via ito-tts / PyTorch / the host chip engine—MCU silicon claims need a careful read (below).
Pro Tip
Use Ito as a host-first prototype this week: get the gated weights, listen to voices D/G at 24 kHz, and wire streaming chunk playback on your laptop. Treat ESP32-S3 timing as a design target, not a measured board result, until Lokutor (or you) publish physical silicon numbers.
What Shipped
FactDetail
Model
Ito streaming TTS (4.4M params, **4.9 MB** int8)
Vendor
Lokutor — Daniel Varela
Voices
2 — female D, male G
Sample rate
24 kHz
Streaming shape
~25 ms first chunk, then ~100 ms chunks (vendor estimates)
Language
English only
Target MCU
ESP32-S3 ($5-class, no NPU)
Code license
GPLv3 (GitHub)
Weights
CC BY-NC-SA + terms (gated on HF)
Quality signal
UTMOS22 ~4.44 vs other MCU TTS (vendor-reported)
Sibling
Oído ASR on the same chip class (separate project)
The blog’s quality comparison leans on UTMOS22 ~4.44 against other MCU-oriented TTS options—useful as a relative score, not a MOS listening study you can cite as independent.
The Silicon Caveat (Read This Twice)
Lokutor is explicit about the verification surface today:
- Bit-exact behavior is demonstrated in Espressif QEMU and on host
- No physical ESP32-S3 board timings have been published yet
- TTFA / RTF figures are estimates, not measured silicon real-time proof
So: do not ship marketing copy that says “Ito runs real-time on ESP32-S3” as a measured fact. Say “designed for ESP32-S3; verified bit-exact in QEMU + host; board timings pending.” That honesty is the difference between a credible MCU audio scoop and a hype trap.
Why MCU Audio People Care
Most “edge TTS” demos quietly assume an NPU phone SoC or a Linux SBC. Ito’s bet is narrower and more interesting: streaming neural speech on a five-dollar ESP32-S3 without an NPU, with a footprint under 5 MB int8 and a chunked playback story (first audio fast, then ~100 ms slices).
If you already watch Lokutor’s lane, pair this mentally with sibling Oído ASR on the same chip class—same vendor story of tiny speech I/O on cheap silicon—without collapsing both into one product. This article stays on Ito TTS.
How to Try It Today (Host Path)
- Read the Oct 2 Lokutor blog
- Request access to gated weights on
lokutor-ai/ito(respect CC BY-NC-SA + terms) - Clone
lokutor-ai/ito(GPLv3) and follow the host /ito-tts/ PyTorch path - Stream voice D or G at 24 kHz; confirm chunk cadence on your machine
- Only then plan an ESP32-S3 bring-up—and measure TTFA/RTF yourself on hardware
Expect English-only prompts and a two-voice palette. Non-commercial weight terms mean product shipping needs a hard license read before you embed.
Conclusion
Ito is a compelling tiny streaming TTS drop for the MCU crowd: ~4.4M / ~4.9 MB int8, two English voices, 24 kHz, GPLv3 code + gated CC BY-NC-SA weights, aimed at a $5 ESP32-S3. Run it on host / QEMU now; keep physical board timings marked unmeasured until someone publishes them. Start from the Lokutor blog and lokutor-ai/ito.
—Anabel ♡
