Introduction

A new reward-tuning method for flow image models just landed with weights you can load right now. MEND ("RL for Flow Models via Proximal Velocity Matching") comes from Shreshth Saini and Alan Bovik at the University of Texas at Austin with Neil Birkbeck, Yilin Wang and Balu Adsumilli at Google. The paper went up on arXiv on October 5, and on October 7 (MYT) the team published nine open LoRA adapters on Hugging Face for Z-Image-Turbo, Stable Diffusion 3.5 Medium and SD3 Medium, along with the full Apache-2.0 training and evaluation code on GitHub.

The headline: MEND gets its gains in about 100 updates, where Flow-GRPO's comparable PickScore adapter took roughly 4,000. You can drop the adapters into diffusers today; base weights download separately.

How MEND works

Most reward post-training either reweights the model's own samples under a KL penalty against a frozen reference, or backpropagates the reward and nudges every sample whether the nudge is worth it. MEND takes a pickier route:

  • Caps rewards per prompt group, so samples that already score well don't get moved at all
  • Proposes moves along the reward gradient for the rest, and accepts one only when its capped reward gain beats a quadratic "displacement price"
  • Regresses onto the resulting velocity targets, with no KL term, no frozen reference model in the loss and no advantage weights
  • Uses an EMA behavior policy during training

The upshot is fewer, more deliberate edits to the model, which is why so few updates are needed. The authors say it applies to any flow backbone with a differentiable reward.

The numbers

These are the paper's own measurements on DrawBench (200 prompts × 5 seeds, 512 px, 40 Euler steps) for SD3.5 Medium:

  • Base SD3.5-M: PickScore 22.35, HPSv2.1 0.280, ImageReward 0.83
  • Flow-GRPO (PickScore, ~4k updates): PickScore 23.52, HPSv2.1 0.316, ImageReward 1.27
  • MEND (PickScore, 100 updates): PickScore 23.70, HPSv2.1 0.319, ImageReward 1.32, at the same DreamSim distance from base images as Flow-GRPO (0.313)
  • MEND three-reward run (300 updates): beats the five-reward DiffusionNFT model on all three rewards it trained on (PickScore, HPSv2.1, CLIPScore)
  • Equal-budget test: MEND reaches PickScore 24.03 versus 23.92 for ReFL and 23.43 for DiffusionNFT
  • Cost: the 100-update SD3.5-M run took 10 GPU-hours on three GB200s

Flow-GRPO still edges it on the aesthetic score, and every main-table MEND config is a single training seed, so treat these as the authors' reported results rather than an independent benchmark.

The nine adapters

Every release includes the original PEFT adapter, an equivalent diffusers LoRA (pytorch_lora_weights.safetensors), checksums and sampling settings. All are rank 32.

  • Z-Image-Turbo: MEND-Z-Image-Turbo-PickScore and MEND-Z-Image-Turbo-HPSv2.1 (100 updates; 1024 px, 9 steps, guidance 0)
  • SD3.5 Medium: PickScore (guidance 4.5), ThreeReward (300 updates), HPSv2.1, CLIPScore and ImageReward (guidance 1.0; 512 px, 40 steps)
  • SD3 Medium: two PickScore seeds (Seed1 and Seed2)

The Z-Image-Turbo pair is the most fun for everyday makers, since that base is Apache-2.0 and already fast at nine steps. The SD3 and SD3.5 adapters ride on Stability's own base-model terms, which you'll need to accept on Hugging Face.

Try it in diffusers

Here's the Z-Image-Turbo snippet from the official model card. Run it in BF16; the authors note FP16 overflows in their pipeline.

import torch
from diffusers import ZImagePipeline

pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo", torch_dtype=torch.bfloat16,
)
pipe.load_lora_weights(
    "shreshthsaini/MEND-Z-Image-Turbo-PickScore",
    weight_name="pytorch_lora_weights.safetensors",
)
pipe.enable_model_cpu_offload()
image = pipe(
    "a small blue book on a large red book",
    height=1024, width=1024, num_inference_steps=9, guidance_scale=0.0,
    generator=torch.Generator(device="cpu").manual_seed(0),
).images[0]
image.save("mend.png")

Swap in MEND-Z-Image-Turbo-HPSv2.1 for the HPS-tuned flavor. The card is upfront that direct diffusers output can differ from the paper's evaluation sampler, and no minimum VRAM is claimed. There's no ComfyUI workflow from the team yet.

Train your own

The GitHub repo ships MEND trainers for SD3.5-M and Z-Image-Turbo, baseline trainers (Flow-GRPO, DiffusionNFT, ReFL, DiffusionOPSD), a six-reward evaluation suite and CPU tests. Training needs Linux and CUDA GPUs; the validated stack is Python 3.11 and PyTorch 2.6+.

git clone https://github.com/shreshthsaini/MEND-RL.git
cd MEND-RL
uv venv --python 3.11 .venv
source .venv/bin/activate
uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
uv pip install -e ".[rewards,dev]"

Why it matters

Reward tuning has usually meant thousands of updates and a frozen reference model eating VRAM. If MEND's 100-update recipe holds up outside the paper, polishing your own flow model toward a reward you care about gets a lot cheaper, and the open trainer means LoRA tinkerers can test that claim this week.

Original Source

Conclusion

MEND is a tidy idea with receipts: move only the samples worth moving, skip the KL babysitter, and ship nine adapters on day one. Load the Z-Image-Turbo PickScore LoRA, run the same seed with and without it, and see if your eye agrees with the reward model.

—Aurelia ♡