LAION Humaneness Voice Small: Open Acting TTS

Introduction
On September 29, 2026, LAION published Humaneness Voice Small on Hugging Face—an experimental English/German text-to-speech and voice-acting model with open CC BY 4.0 weights, ten stage checkpoints from one continuous S1→S10 run, and a runnable code/infer.py path (CUDA + the separate MOSS Audio Tokenizer v2).
Primary artifact: laion/Humaneness-Voice-Small (card created ~19:39 MYT Sep 29). This is a research release with known prompt sensitivity—not a production-certified studio voice—but it is a same-day open-weight acting stack makers can pull and audition today.
Pro Tip
Start with stage S3 (LAION’s provisional default) or S4 (best held-out validation loss) and a short CAPTION + quoted TRANSCRIPT prompt before you try the heavier structured GENERAL: / SCRIPT: surface. Keep reference clips clean, under about three seconds, and never use the target take as its own reference.
What Landed on the Card
The model card is unusually concrete for a day-0 research drop:
- Open weights under CC BY 4.0, plus training/inference source and run stats in-repo
- Full state about ~715 million parameters; semantic backbone initialized from pretrained Qwen3-0.6B; ~112.8M Talker + bridge trained fresh
- Codec path via separately loaded MOSS Audio Tokenizer v2; 48 kHz reconstruction; 12 RVQ codebook indices per 80 ms frame (150 indices/sec)
- Ten BF16 model-only checkpoints (
checkpoints/S1…S10) from one continuous ladder run on 32 JUPITER Booster nodes × 4 GPUs - Prompted voice acting: emotion, timing, bursts, ambient/sound-event fields, optional reference audio
- Languages called out: English and German
- Not
AutoModel.from_pretrained()— use the repo’scode/infer.pyand read the inference notes first
Closed commercial TTS (Eleven v4, Cartesia, Fish) still win for agents and polish. Humaneness Voice Small is the open ladder for people who want to study and steer expressive delivery locally.
Prompt Surfaces That Actually Trained
LAION documents the outer <user_inst> template with fields like Reference(s), Instruction, Tokens (codec-frame budget), Quality, Sound Event, Ambient Sound, Language, and Text. Two practical entry points:
Caption + exact transcript (everyday readable speech):
CAPTION: Warm, softly amused, conversational.TRANSCRIPT: "I cannot believe we made it."
**Structured GENERAL / SCRIPT** (when delivery, durations, or bursts matter):GENERAL: frightened, breath-held, intimate, increasingly panicked, adult, low-registerSCRIPT:(quiet alarm, searching hesitation) [3.7 seconds duration] I locked the back door... didn't I?(fear rising, clipped self-correction) [4.2 seconds duration] I heard the latch click, but now it is standing open.(whispered panic, sharp inhale) [5.1 seconds duration] Wait—please don't move; there is someone breathing in the hallway.The card is honest that structured C01 timing/burst control is still shaky (higher WER than caption modes). Treat durations as training variations, not hard guarantees.
## How to Run It Today
1. Open the primary card: https://huggingface.co/laion/Humaneness-Voice-Small
2. Pull the repo (checkpoints + `code/`) and accept the separate **MOSS Audio Tokenizer v2** download
3. Use a CUDA environment with PyTorch, Transformers, NumPy, SoundFile, Torchaudio, and Safetensors
4. Smoke-test with the shipped example (stage S3, 75 frames ≈ six seconds of speech budget):python code/infer.py --stage S3 --language en --frames 75 \ --prompt $'CAPTION: Warm, softly amused, conversational.\nTRANSCRIPT: "I cannot believe we made it."' \ --text 'I cannot believe we made it.' --output example.wavAdd --reference-wav reference.wav for a distinct clean clip under ~3 s when you need identity. Prefer the public listening Space on the card when you want audited takes without wiring the full stack.
Caveats Worth Keeping
- Research release: disclose synthetic speech; avoid impersonation or deceptive use
- Attribute LAION / Humaneness Voice Small under CC BY 4.0; respect Qwen3 and MOSS tokenizer licenses upstream
- Acting metrics on the card are proxy scores and automatic rewards, not a human MOS leaderboard win
- Zero download/like heat at discovery—this is a primary artifact story, not a trending amplification piece
Conclusion
Humaneness Voice Small is the clean Sep 29 open-weight audio drop in-lane: a LAION EN/DE voice-acting TTS with ten ladder checkpoints, a documented prompt grammar, and a same-day infer.py path beside MOSS Audio Tokenizer v2. Grab the Hugging Face card, audition S3 vs S4 on a familiar script, and keep Eleven v4 / Cartesia / Fish in their closed-API lanes when you need production agents instead of research acting control.
—Anabel ♡
