LAION Humaneness Voice Converter: Open Chatterbox Voice Conversion That Keeps the Emotion

Introduction
Humaneness Voice Converter is a new open voice-conversion model from LAION, the nonprofit behind the LAION datasets. It went up on Hugging Face on October 11, 2026. You give it a spoken recording and a short clip of a different, consenting speaker, and it says the same words in that second voice while keeping more of the original performance: the timing, the emphasis, and the emotion.
It's built on Resemble AI's Chatterbox, and it's open weights. LAION's additions are licensed CC BY 4.0, and the bundled Chatterbox parts keep their upstream MIT license. You can listen to it today in the free comparison demo Space or run it on your own GPU with the included infer.py.
This is a speech-to-speech converter, not a text-to-speech model. It's a sibling to LAION's Humaneness Voice Small acting TTS from late September.
Pro Tip
Use a clean 5–10 second target clip with one speaker. That's the range LAION tested, and both inputs should be 1–30 seconds of mono audio. If a take sounds off, try --candidates 8, which generates eight versions and keeps the best-scoring one.
What Shipped
From the Humaneness Voice Converter model card:
- A complete checkpoint, not a pile of LoRAs. It bundles the Chatterbox speech-token extractor, the CAMPPlus speaker encoder, the new score-conditioned flow decoder, the HiFT vocoder, and LAION's frozen Humaneness Ears Medium scoring model.
- Emotion-aware conditioning. Ears Medium predicts 99 scores per clip (40 emotions, 57 voice attributes, plus genuineness and vocal-burst blending). Style and emotion come from your source recording, and stable traits like timbre, age, and gender come from the target.
- Output: a 24 kHz WAV, generated in 10 flow steps.
- Languages: English and German were in the training and evaluation data. LAION says other languages may work but makes no quality promise.
- Training: LAION trained all 112.6 million flow-decoder parameters with a group-relative reinforcement learning reward that weighted how enjoyable the audio sounds (40%), how closely it keeps the source's top emotions (30%), and how much it sounds like the target speaker (30%).
The Numbers
LAION scored 40 curated source clips against two target voices. These are automatic proxy scores, not human listening tests, and LAION says there's no independent human MOS or word-error result yet.
DecoderContent Enjoyment ↑Production Quality ↑Emotion error ↓Target speaker match ↑
Original Chatterbox
6.169
7.766
0.113
0.821
Humaneness VC (default)
6.453
8.009
0.092
0.773
Humaneness VC, best of 16
6.460
8.009
0.072
0.795
The trade-off is honest and worth knowing. The new model sounds more polished and keeps the emotion closer to the original, but it sounds a little less like the target speaker than plain Chatterbox. Picking the best of several takes wins some of that back.
It's quick, too. On one NVIDIA GH200, a 6.5-second clip took about 0.47 seconds for one take in FP32, and about 0.67 seconds for 16 takes at once in BF16 (see SPEED.md). Those timings don't include model loading or ranking.
How to Try It Today
- Just listen: open the demo Space. It plays the same 40 sources through original Chatterbox, Humaneness VC, best-of-N picks, and other voice converters side by side, including German clips.
- Run it locally with a CUDA build of PyTorch:
python -m pip install -r requirements.txt
python -m pip install --no-deps chatterbox-tts==0.1.7
hf download laion/Humaneness-Voice-Converter --local-dir humaneness-vc
python humaneness-vc/infer.py --model-dir humaneness-vc \
--source source_speech.wav --target consenting_target_voice.wav \
--output converted.wav --seed 20261011 --flow-steps 10- Rank a few takes: add
--candidates 8 --microbatch 8 --precision bf16 --metrics selected.metrics.jsonto generate eight versions and keep the top one.
LAION tested this with Python 3.13, PyTorch 2.9.1, and Chatterbox 0.1.7 on its own cluster, so check your versions before you install.
Why It Matters
Most voice converters either keep the target voice well and flatten the acting, or keep the acting and drift off the voice. LAION is openly tuning for the performance side and publishing the trade-off numbers instead of hiding them. For dubbing tests, character voices, game dialogue prototypes, and building speech datasets, that's a useful open option you can actually inspect and retrain. The model card is clear that it's a research release: only use voices you have permission to use, label converted speech, and don't use it to impersonate anyone.
Conclusion
Humaneness Voice Converter is LAION's open, CC BY 4.0 voice-conversion checkpoint built on Resemble AI's Chatterbox. It turns one recording into another speaker's voice from a 5–10 second target clip, outputs 24 kHz audio, and scores higher on predicted quality and emotion match than original Chatterbox, at a small cost in speaker similarity. You can hear it now in the Hugging Face demo or run it with infer.py on a CUDA GPU.
Primary: Humaneness Voice Converter model card (Oct 11, 2026), Humaneness Voice Converter Demo Space, measured speed notes, Resemble AI Chatterbox.
—Anabel ♡
