Introduction

Amazon made Nova 2.5 Sonic generally available on Oct 5, 2026. It's a speech-to-speech model for real-time voice agents: audio goes in, expressive audio comes out, with no separate speech-to-text and text-to-speech steps in between. It's closed weights, and you run it today through Amazon Bedrock in four AWS Regions. Primary source: the AWS announcement Announcing Amazon Nova 2.5 Sonic.

The upgrade is mostly about the brain behind the voice. AWS says reasoning, instruction following, and tool-calling accuracy all improve, and latency drops. On the same day, the open-source Strands Agents team shipped Bidi Agents as GA, giving you a few-line Python path to a live mic-and-speaker agent on the new model.

Pro Tip

If you already run Nova 2 Sonic, this is a cheap test. AWS says pricing is the same as Nova 2 Sonic, so try your existing prompts and tools on 2.5 and listen for the difference on multi-step tasks, like "look up my order, check if it's returnable, then start the return."

What Shipped

FactDetail

Model

Amazon Nova 2.5 Sonic

Vendor

Amazon (AWS)

Modality

Speech-to-speech (voice and text in the same session)

Status

Generally available, Oct 5, 2026

Weights

Closed, API only

Where to run

Amazon Bedrock: US East (N. Virginia), US West (Oregon), Europe (Stockholm), Asia Pacific (Tokyo)

Context window

256K

Voices

Expressive voices in seven languages

Agent features

Asynchronous tool calling, controllable turn-taking

Session limit

8 minutes per connection (per the Strands team)

Price

Same as Nova 2 Sonic

Strands' Nova Sonic docs list the language set as English (US, UK, AU, IN), French, Italian, German, Spanish (US), Portuguese (BR), and Hindi, with both masculine- and feminine-sounding voices.

What's Actually Better

AWS frames 2.5 as the version that can hold a real conversation and get work done in it:

  • Reasoning and instruction following: the agent keeps context across a call and follows multi-step instructions
  • Tool calling: more accurate calls to the tools you connect, and they can run asynchronously while the conversation keeps going
  • Latency: lower than before, so turns feel quicker and more natural
  • Mixed input: voice and text can share one session

That matters because speech-to-speech models pick up how someone talks, not just what they say. The Strands team put it nicely: a voice model can hear that a caller sounds stressed without the transcript having to say "I'm stressed."

AWS didn't publish new latency numbers or benchmark scores in the announcement, so treat "lower latency" as a vendor claim until you time it on your own calls.

Try It in a Few Lines (Strands Bidi Agents)

The Bidi Agents GA post (Oct 5) shows a working voice loop. Install the SDK with voice and local audio extras (you'll also need PortAudio and AWS credentials with Bedrock access):

pip install "strands-agents[bidi-all,bidi-pyaudio]"

Then run a mic-to-speaker agent:

import asyncio
from strands.bidi.agent import BidiAgent
from strands.bidi.io import AudioIO
from strands.bidi.models import BedrockNovaSonicModel

async def main():
    model = BedrockNovaSonicModel(model_id="amazon.nova-2-5-sonic")
    agent = BidiAgent(
        model=model,
        system_prompt=(
            "You are a helpful voice assistant. "
            "Keep responses concise and conversational."
        ),
    )
    audio = AudioIO(audio_processor=True)
    # Runs until you press Ctrl+C.
    await agent.run(inputs=[audio.input()], outputs=[audio.output()])

asyncio.run(main())

The model ID above is the one in the Strands post. If Bedrock rejects it in your Region, copy the exact ID from the Bedrock console.

Two Bidi Agents features help in practice. Echo suppression stops the agent from interrupting itself when your speakers are on. Auto-reconnect quietly restarts the connection at a turn boundary before Nova 2.5 Sonic's 8-minute cap, so long calls keep going with context carried over. Bidi Agents can also swap in OpenAI or Gemini realtime models without rewriting your tools.

Who Should Care

  • Voice-agent builders on AWS: a free upgrade path at the same price, with better tool use
  • Creators prototyping talking characters or guides: a hosted voice that can look things up mid-conversation
  • Teams comparing speech-to-speech stacks: a fresh contender next to GPT-Realtime and Gemini Live, now behind one SDK

It's not for you if you need open weights, on-device inference, or languages outside the seven listed.

Conclusion

Nova 2.5 Sonic is Amazon's GA speech-to-speech upgrade: smarter reasoning, more accurate async tool calls, lower latency, a 256K context, and expressive voices in seven languages, on Bedrock in four Regions at Nova 2 Sonic pricing. Start with the AWS announcement, then wire up a test call with Strands Bidi Agents.

—Anabel ♡