Fizgig v6.5: Train Qwen-Image 2.1 LoRAs

Introduction
Qwen-Image 2.1 already paints sharp. The pain starts when you train a LoRA on it: partway through a run the image can collapse into texture or wobble, and a falling loss number will not warn you.
Fizgig v6.5.0 from ShootTheSound makes Qwen-Image 2.1 a first-class base inside the studio — LoRA and LoKR, the same workbench tools Flux 2 Klein / Krea 2 / MiniMax H3 already enjoy — and ships Fizgig’s own frozen training adapter so those collapses stop happening. Saved adapters drop into ComfyUI’s standard LoRA loader.
What Shipped in v6.5
Per the v6.5.0 release notes (published 2026-09-27 UTC):
- Base model: Qwen Image 2.1 on the Training tab (experimental), with LoRA or LoKR, Adaptive LR or Automagic, EMA, Context LoRA, pause/resume.
- Turbo previews: Viggle’s 6-step turbo LoRA during training (the author notes it can glitch — try Turbo strength
0and ~25 steps for cleaner samples). - Per-image loss watch: problem-image detection, per-image LR, look-outlier warm-up, auto-recaption via the same Qwen3-VL captioner the Captions tab uses.
- Workbench: Repair Studio, LoRA the Explorer, Profiler, Extract, LoRA Royale — all on Qwen.
- Driver system: Qwen-Image 2.1 is the first model added through Fizgig’s new driver interface (files + LoRA format + model code behind one API). A PR guide for adding more models is promised soon.
- Coming soon: edit training for Qwen Image 2.1.
Fizgig is a standalone trainer, not a ComfyUI custom node. It writes kohya-format .safetensors that land in ComfyUI/models/loras/.
The Training Adapter (The Real Fix)
Qwen 2.1 LoRAs have a habit of collapsing into texture mid-run. Fizgig’s fix is a small training adapter that:
- stays frozen and active while you train
- switches off for previews
- is stripped from the saved LoRA, so the result still runs on the plain base
It is on by default in every Qwen preset, downloaded with the model pack, and also published at ShootTheSound/Fizgig-Qwen-Image-2.1-Training-Adapter for other trainers. The author trained it at higher resolution than the prior “training assistant” cards and reports sharper LoRAs plus the collapse fix.
Three Presets at 0.5 MP
All three presets use adamw8bit, EMA 0.98, the training adapter, 30 epochs, and save every epoch. Target resolution is 0.5 megapixels — not 512². That is about 704×704 square; Fizgig buckets other shapes at the same pixel count (e.g. 624×784 at 4:5, 928×528 at 16:9).
PresetRankLearning rateBest for
Qwen 2.1 Fast (default)
8
Adaptive 2e-4–4e-4
Most subjects; quickest likeness; holds skin detail
Qwen 2.1 Standard
16
Adaptive 1e-4–2e-4
Bigger / mixed datasets
Qwen 2.1 Style
16
Flat 1.5e-4
Styles (adaptive LR tends to climb and overbake)
Why train “fast”? Qwen is natively sharp; a long run on softer photos slowly trades that sharpness for your dataset’s texture. Shorter, 0.5 MP runs keep more of the base look. Scrub epochs in LoRA Royale if a later save softens.
Train on ~10 GB Cards
10 GB is the floor for Qwen Image 2.1 in Fizgig. Base precision Auto picks bf16, INT8, or 4-bit NF4 from free VRAM and resolution, then sizes Blocks Swap to match (quantize before swap). On cards that cannot hold the full text encoder, Fizgig loads an 8-bit encoder automatically.
Author stress tests on an RTX 5090 capped at 12 GB: INT8, ~9.9 GB peak with previews. Capped at 10 GB: NF4, ~7.3 GB peak. Real 10–12 GB cards make the same choices but run slower than the 5090’s ~0.9–1 s/step.
Get It Running
- Grab the v6.5.0 release (or latest from the repo).
- Preferences → Download models for me in the Qwen Image 2.1 section (~34 GB: DiT, VAE, text encoder, training adapter, turbo LoRA, plus Qwen3-VL captioner/tokenizer if missing).
- Training tab → Base Model = Qwen Image 2.1 → load a preset → train.
- Drop the saved
.safetensorsintoComfyUI/models/loras/and load with the standard LoRA node.
Original Source
https://github.com/shootthesound/Fizgig/releases/tag/v6.5.0
Secondary writeup: https://comfyui-wiki.com/en/news/2026-09-27-fizgig-v6-5-qwen-image-2-1
Training adapter: https://huggingface.co/ShootTheSound/Fizgig-Qwen-Image-2.1-Training-Adapter
Conclusion
Fizgig v6.5 is the missing trainer lane for Qwen-Image 2.1: a real base-model slot, a frozen adapter that kills texture collapse, sensible 0.5 MP presets, and Comfy-ready LoRAs that still fit on a 10 GB card. If you already live in that Qwen stack, this is the cleanest path from dataset to loader.
—Aurelia ♡
