Introduction

On October 1, 2026, Suno shipped Speech (beta) — a closed generative audio model the company calls the first to create spoken voice and its soundtrack together in one take. Primary sources: Introducing Speech (beta) by Chief Product Officer Jack Brody, and the same-day release note (Android / iOS / web / Create).

Closed / app-only — no open weights. You can run it today inside Suno on mobile and web after a month of small-group testing; the beta is now open to everyone.

Pro Tip

Prompt Speech like a mini production brief, not a song lyric sheet: lead with the spoken content (poem, pep talk, bedtime story), then pin voice + music in a second clause — e.g. “warm British storyteller over soft piano” vs “stadium hype over drums.” Keep takes short while accents still wander in beta.

What Shipped

Concrete facts from Suno’s Oct 1 blog + release note:

FactDetail

Model / feature

Speech (beta) — spoken audio + original background music as one cohesive track

Vendor

Suno

Weights

Closed (Suno product surface)

Where to run

Suno mobile and web (release note tags: Android, iOS, web, Create)

Rollout

~1 month closed test → open beta to all users on Oct 1, 2026

Example uses (vendor)

Bedtime stories over soft piano; hype speeches over stadium drums; ASMR grocery lists

Suno frames Speech under “creative entertainment”: type an idea, poem, or written piece, then describe the voice and musical style. The company is explicit that beta means beta — British accents can drift toward Australian and back, and dramatic pauses can land very dramatic.

This is a different product surface than Suno v6 (Sep 9 music models with Warner / BMG / Believe) or Voices (record-once singing identity). Speech targets spoken-word + score in a single generation, not a sung track.

Why Audio Makers Care

Most TTS stacks give you a dry read; most music models give you a song. Speech collapses the two into one consumer loop — useful for storytime, meditations, short-form narration, and joke “overproduced” voice notes without bouncing stems between a TTS API and a music API. For makers already in Suno, it is another Create canvas beside v6 music, not a replacement ASR or agent TTS lane (keep Eleven / Cartesia / MAI there).

How to Try It

  1. Read Introducing Speech (beta) and skim the release note
  2. Open Suno on web or mobile → Create → pick Speech (beta)
  3. Paste spoken text, then describe voice + musical bed in the same prompt
  4. Export a few variants; A/B accent stability and pause timing before you ship to clients

A two-minute bedtime story about a fox who learns to share fireflies.Voice: soft, warm storyteller, gentle smile in the tone.Music: quiet acoustic piano and soft night pads, lullaby tempo, no lyrics.

Conclusion

Oct 1’s Suno Speech (beta) is the clearest consumer generative-audio drop in this scout window: closed weights, runnable same day on mobile and web, and a concrete claim — speech + soundtrack in one take. Start from the official blog, keep expectations beta-shaped, and leave agent TTS / streaming ASR comparisons to the Microsoft MAI and Eleven lanes.

—Anabel ♡