Introduction

HiDream just walked into the big leagues of AI video. HiDream-O1-Video-1.0, the closed video model from HiDream.ai, debuted at #6 on the Artificial Analysis Image to Video with Audio leaderboard on October 9. It turns a single still into a 1080p clip of 5 to 20 seconds with natively synced sound, and you can run it today through the HiHarness API or inside vivago R1 Studio.

The arena placement is the fresh news here. The model quietly landed in vivago R1 Studio on September 30 next to HiDream-O1 Image 2.0 and Editing 1.5, but this is the first independent, blind-vote read on how it stacks up against the heavy hitters.

What shipped

  • Model: HiDream-O1-Video-1.0 (HiDream.ai)
  • Modality: image-to-video with native audio
  • Weights: closed, API and app only
  • Output: 1080p, 5 to 20 seconds, with the length chosen by the scene rather than fixed up front
  • Price: $5.80 per minute of 1080p video with audio on HiHarness, roughly $0.10 per second
  • Where to run it: HiHarness API and vivago R1 Studio

HiDream pitches it as a "native omnimodal" video model built for physical consistency, meaning motion, objects and sound are generated together instead of bolting audio on afterward.

How it ranks

In the Artificial Analysis Video Arena, HiDream-O1-Video-1.0 sits just behind Dreamina Seedance 2.0 (720p) and lands in a tight cluster with MiniMax H3, Vidu Q4 Preview and Gemini Omni Flash. It finishes ahead of Wan 3.0.

That's a strong first showing for a lab most creators know from open image weights. Landing in the same pack as H3 at about ten cents a second makes it a real option for audio-on clips, not just a curiosity.

The prompts Artificial Analysis used

The arena thread showed the model on three image-to-video prompts. These two are short, complete, and handy templates for testing motion, sound and physics on any I2V model:

Deadlift PR attempt, bar bends, pulls slow, locks out, screams, drops it. Gripping, pulling, bar bending, locking out, screaming, dropping.

Passengers flip newspapers open in sync, pages rustle together, train rocks gently. Papers unfolding, rustling, train hum.

Notice the pattern: a short action line, then a second line listing the motions and sounds you want to hear. That two-part shape is a neat trick for any audio-capable video model.

Run it on HiHarness

The current HiHarness docs expose a narrower surface than the launch pitch. Today's endpoint is single image in, one video out, with an optional prompt. A few details worth knowing before you wire it up:

  • One reference image, as a public URL or raw Base64, between 40 KB and 20 MB
  • Output keeps your input image's aspect ratio
  • Duration is dynamic by default, or set force_10s to lock it at 10 seconds
  • Jobs are async: submit, grab task_id, then poll for the result
  • No request fields are documented yet for reference video, a resolution picker, or audio control

Submit a job:

curl --request POST 'https://hiharness.hidreamai.com/api/maas/gw/v1/videos/generations' \
  --header 'Authorization: Bearer <YOUR_API_KEY>' \
  --header 'Content-Type: application/json' \
  --data-raw '{
    "model_id": "HiDream-O1-Video-1.0",
    "image": "https://example.com/your-still.png",
    "prompt": "The subject turns toward the camera, with gentle cinematic motion"
  }'

Then poll with the returned ID:

curl --get 'https://hiharness.hidreamai.com/api/maas/gw/v1/videos/generations/results' \
  --header 'Authorization: Bearer <YOUR_API_KEY>' \
  --data-urlencode 'task_id=<TASK_ID>'

One gotcha: result.status of 1 only means every subtask has finished, not that it worked. Check each sub_task_results[].task_status. A 1 there means success and comes with the video URL, 3 means generation failed, and 4 means the clip didn't pass safety review.

Original Source

Artificial Analysis leaderboard announcement:

API docs: HiHarness HiDream-O video models

Conclusion

HiDream-O1-Video-1.0 earns its seat at the table: 1080p, sound baked in, clip length that follows the scene, and a price that won't scare off a weekend experiment. The public API is still image-only, so treat the bigger omnimodal promises as "coming" until the docs catch up. For now, grab your favorite still, write one action line and one sound line, and let it roll.

—Aurelia ♡