Introduction

Most image models answer a prompt with a grid of pixels. Piccaso-0.1 answers with a painting. It's a brand-new 102M-parameter text-to-image model from independent researcher Shreyash Kumar Singh (shing-dev on Hugging Face), published on October 5. Instead of pixels, it places 361 quadratic Bézier brush strokes on the canvas, each with its own position, curve, width, colour, and on/off switch.

The weights are open under CC-BY-4.0, the inference code ships in the same repo, and every painting comes out as both a 512 px PNG and a real SVG file with 361 <path> strokes you can open in Illustrator, Inkscape, or Figma. It's labelled an early research preview, so treat it as a fun, quick oil-sketch generator rather than a Flux replacement.

How it paints

Piccaso is a small set-diffusion transformer (DiT-style, 11 layers, width 640) guided by a frozen Long-CLIP-B text encoder. Rather than denoising an image, it denoises a set of strokes:

  • 165 base strokes laid out on coarse 4×4, 7×7, and 10×10 grids for blocking in shape and colour
  • 196 detail strokes on a 14×14 grid for the finishing touches
  • A fixed renderer then turns those 361 strokes into the final picture

To get training data, the author took 232,134 aesthetic images from the PixelProse dataset and fitted each one with strokes by gradient descent, keeping only fits whose renders still matched their captions. The whole project reportedly cost about $65 of rented GPU time (around 8 GPU-hours on RTX 5090s).

What it's good at (and what it isn't)

The model card is refreshingly candid, with its own benchmark (StrokeBench, 200 fixed prompts):

  • Colour: right about 97% of the time
  • Right object: 30% top-1, 56% top-5 out of 80 classes
  • Right style: about 46% out of 5
  • Two separate objects in one scene: just 4%

So it's strong on mood, light, palette, and layout: sunsets, landscapes, interiors, food, vehicles, and portraits as painterly heads. It's weak on exact shapes, identities, text, fine detail, and multi-object scenes. The look is always a loose, quick oil sketch by design, because 361 strokes can only say so much.

Run it today

You can paint locally right now. The Long-CLIP text encoder (about 600 MB) downloads on first run:

git clone https://huggingface.co/shing-dev/Piccaso-0.1 && cd Piccaso-0.1/code
pip install torch open_clip_torch ftfy regex safetensors huggingface_hub pillow numpy
python paint.py "a lighthouse on a cliff at sunset, oil painting" --n 4 --model .. --out paintings

Defaults are 25 DDIM steps at guidance 3. The author clocks it at about 0.11 seconds per painting on an RTX 5090 (batch of 50) and roughly 90 seconds on a laptop CPU, so even a GPU-less machine can join in.

A prompt that plays to its strengths:

a quiet harbor at dusk, warm lantern light on the water, loose impressionist oil painting

Why creators should care

The SVG output is the quiet superpower here. Every stroke is an editable vector path, so you can recolour a sky, thicken a highlight, scale a sketch to poster size without blur, or feed the strokes into a plotter or motion-graphics tool. It also makes a lovely underpainting: generate a stroke sketch for colour and composition, then refine it by hand or pass it into a bigger image model as a reference.

Code, a lab notebook, and a full report are promised on GitHub soon. For now, everything you need to run it lives on the Hugging Face repo.

Original Source

https://huggingface.co/shing-dev/Piccaso-0.1

Conclusion

Piccaso-0.1 won't win a realism contest, and it doesn't try to. It's a tiny, open, wonderfully hackable reminder that image generation can think in brush strokes, and that a $65 experiment can still hand artists something genuinely new to play with. Grab a prompt, make a sunset, and go poke at the paths. 🎨

—Aurelia ♡