Introduction

Most video models get tired. Push an autoregressive generator past its training window and faces melt, rooms rearrange themselves, and the colors slowly wash out. SGF+ (Self Gradient Forcing Plus) goes after that problem directly. It's a new open text-to-video method from Tsinghua University and JD's Joy Future Academy (with CUHK). It learns from 5-second rollouts, then keeps streaming one continuous video for up to 24 hours with no long-video fine-tuning at all.

Paper, code, and two checkpoints went public on October 7 under Apache-2.0. You can run it locally today from the GitHub repo with weights from Hugging Face. On October 8 it was also climbing HF Daily Papers.

What Shipped

  • Paper: arXiv 2610.10429, submitted October 7, 2026
  • Weights: chunkwise/model.pt and framewise/model.pt, about 11.4 GB each, plus a diagnostic SGF checkpoint
  • Code: inference scripts, training launchers, and the gradient-conflict experiments
  • Base: a Wan2.1-T2V-1.3B student distilled from the Wan2.1-T2V-14B teacher, starting from Causal Forcing weights, with 4 denoising steps
  • License: Apache-2.0 for code and checkpoints
  • Project page: side-by-side videos plus a full 24-hour rollout

The Trick: Give Memory Its Own Brain

An autoregressive video model does two jobs at once. It denoises the frames it's making right now, and it writes context: the key-value memory that every later frame leans on.

The earlier Self Gradient Forcing (SGF) let the future train that memory-writing job. When the team looked closely, though, they found the two jobs pulling in opposite directions. Across 128 prompts and four timesteps, all 512 paired gradient similarities came out negative, so on shared weights the updates partly cancel each other.

SGF+ fixes this by giving each job its own full set of parameters. Causal attention still connects them, and training uses the same objective with no extra losses, no extra video data, and no longer training horizon. The ablations are telling: splitting only the FFN layers (MoE-style) or only attention falls short. The win comes from separating everything.

The Numbers

From the paper's 240-second chunkwise test on MovieGen prompts:

  • Subject consistency: 98.21 (Self Forcing 94.94, SGF 97.72)
  • Background consistency: 97.31 (SF 95.21)
  • Temporal flicker score: 97.57 (SF 93.54)
  • Aesthetic quality: 64.74 (SF 55.12, SGF 62.54)
  • Imaging quality: 71.39 (SF 68.38)

One honest caveat: the Dynamic Degree score drops (56.92 vs 93.90 for Self Forcing). The authors point out that scene jumps and deformation inflate that metric. Expect steadier long shots, not frantic action.

Cost stays modest. Inference memory rises from 24.85 GB to 27.96 GB, and 81 frames still take about 4.97 seconds on the authors' setup, the same as the baselines.

Run It Today

The default launcher renders 963 latent frames, about 240 seconds at 16 fps. It spreads across 8 GPUs when it sees them and runs serially on one GPU otherwise. Prompts go one per line, the same as the bundled prompts/test_prompt.txt. Save the prompt from the next section as a single line in prompts/my_prompts.txt before the last command.

git clone https://github.com/Zihan-Su/Self_Gradient_Forcing_Plus.git
cd Self_Gradient_Forcing_Plus
conda create -n sgf_plus python=3.10 -y
conda activate sgf_plus
pip install -r requirements.txt
pip install flash-attn --no-build-isolation
python setup.py develop

# Wan2.1 1.3B + 14B, Causal Forcing init, SGF+ checkpoints, prompt list
bash scripts/download_weights.sh

# 4-minute chunkwise rollout with your own prompt file
NUM_OUTPUT_FRAMES=963 SEED=42 OUTPUT_ROOT=outputs/demo \
  bash scripts/infer_sgf_plus.sh chunkwise hf_weights/chunkwise/model.pt prompts/my_prompts.txt

Use chunkwise (3 frames per block) for the smoothest long takes, or swap in framewise with hf_weights/framewise/model.pt for frame-by-frame streaming.

Prompt That Suits Long Rollouts

SGF+ shines on one steady world that slowly evolves, so give it a scene with a natural rhythm:

A cozy lighthouse keeper's cottage window at dusk, rain softly streaking the glass, a warm desk lamp glowing beside a steaming teacup, waves rolling against distant rocks as the lighthouse beam sweeps slowly across the sea, gentle cinematic lighting, calm static medium shot, muted film texture.

Good to Know

  • There's no ComfyUI node or hosted API yet. For now this is a Python and terminal workflow.
  • Output follows Wan2.1-1.3B's 480p-class training shape, so think of it as a long-take engine, not a 4K hero-shot model.
  • Training is heavy (about 98 GB peak memory per GPU in the paper), but inference needs about 28 GB, so a single 32 GB+ card can handle it.
  • Pair it with this morning's GRACE latent-compression story and you can see where Wan2.1 research is heading: longer, faster, and steadier.

Original Source

Co-author Weiyang Jin's announcement:

Conclusion

SGF+ is a nice reminder that sometimes the breakthrough is a cleaner division of labor, not a bigger model. Give memory and drawing their own weights, and a 1.3B model trained on five-second clips can dream through an entire day. Go put a lighthouse on a 24-hour loop and see how long it keeps its glow.

—Aurelia ♡