Introduction

Lokutor (Daniel Varela) published Ito—a tiny streaming neural TTS stack aimed at a $5-class ESP32-S3 with no NPU. Primary sources: the Oct 2, 2026 blog Ito: streaming neural speech on a $5 chip, the gated HF card lokutor-ai/ito (blog also mentions ito-tts-v3), and GitHub lokutor-ai/ito.

Headline numbers: ~4.4M parameters / ~4.9 MB int8, two voices (female D / male G), 24 kHz, streaming with a claimed ~25 ms first audio chunk then ~100 ms chunks. English only. You can run it on a laptop today via ito-tts / PyTorch / the host chip engine—MCU silicon claims need a careful read (below).

Pro Tip

Use Ito as a host-first prototype this week: get the gated weights, listen to voices D/G at 24 kHz, and wire streaming chunk playback on your laptop. Treat ESP32-S3 timing as a design target, not a measured board result, until Lokutor (or you) publish physical silicon numbers.

What Shipped

FactDetail

Model

Ito streaming TTS (4.4M params, **4.9 MB** int8)

Vendor

Lokutor — Daniel Varela

Voices

2 — female D, male G

Sample rate

24 kHz

Streaming shape

~25 ms first chunk, then ~100 ms chunks (vendor estimates)

Language

English only

Target MCU

ESP32-S3 ($5-class, no NPU)

Code license

GPLv3 (GitHub)

Weights

CC BY-NC-SA + terms (gated on HF)

Quality signal

UTMOS22 ~4.44 vs other MCU TTS (vendor-reported)

Sibling

Oído ASR on the same chip class (separate project)

The blog’s quality comparison leans on UTMOS22 ~4.44 against other MCU-oriented TTS options—useful as a relative score, not a MOS listening study you can cite as independent.

The Silicon Caveat (Read This Twice)

Lokutor is explicit about the verification surface today:

  • Bit-exact behavior is demonstrated in Espressif QEMU and on host
  • No physical ESP32-S3 board timings have been published yet
  • TTFA / RTF figures are estimates, not measured silicon real-time proof

So: do not ship marketing copy that says “Ito runs real-time on ESP32-S3” as a measured fact. Say “designed for ESP32-S3; verified bit-exact in QEMU + host; board timings pending.” That honesty is the difference between a credible MCU audio scoop and a hype trap.

Why MCU Audio People Care

Most “edge TTS” demos quietly assume an NPU phone SoC or a Linux SBC. Ito’s bet is narrower and more interesting: streaming neural speech on a five-dollar ESP32-S3 without an NPU, with a footprint under 5 MB int8 and a chunked playback story (first audio fast, then ~100 ms slices).

If you already watch Lokutor’s lane, pair this mentally with sibling Oído ASR on the same chip class—same vendor story of tiny speech I/O on cheap silicon—without collapsing both into one product. This article stays on Ito TTS.

How to Try It Today (Host Path)

  1. Read the Oct 2 Lokutor blog
  2. Request access to gated weights on lokutor-ai/ito (respect CC BY-NC-SA + terms)
  3. Clone lokutor-ai/ito (GPLv3) and follow the host / ito-tts / PyTorch path
  4. Stream voice D or G at 24 kHz; confirm chunk cadence on your machine
  5. Only then plan an ESP32-S3 bring-up—and measure TTFA/RTF yourself on hardware

Expect English-only prompts and a two-voice palette. Non-commercial weight terms mean product shipping needs a hard license read before you embed.

Conclusion

Ito is a compelling tiny streaming TTS drop for the MCU crowd: ~4.4M / ~4.9 MB int8, two English voices, 24 kHz, GPLv3 code + gated CC BY-NC-SA weights, aimed at a $5 ESP32-S3. Run it on host / QEMU now; keep physical board timings marked unmeasured until someone publishes them. Start from the Lokutor blog and lokutor-ai/ito.

—Anabel ♡