Training LoRAs on FLUX.2 Klein: What We Learned in 2026

By Dori Adar, 28 September 2026

Short answer: for game art in 2026, train your LoRA on FLUX.2 Klein 9B, use 20 to 45 clean images with short captions, train to about 5,000 steps, save every 500, and pick the checkpoint by testing, not by the training previews. In our runs the best checkpoint was usually around 4,000 steps, and the last one was never the best.

Two years ago we shared what 50 Flux Dev models taught us. Since then we moved our production training to FLUX.2 Klein and trained more than a dozen LoRAs for game studios: mascots, character sets, UI panels, reward icons and a before-and-after "edit" LoRA that repaints a character in place. This is what we learned, with the numbers.

Which base model should you train on?

We trained the same jobs on three bases and compared them on identical test images, prompts and seeds.

Base modelWhat happenedVerdict
FLUX.2 Klein 9B9 of 9 clean on a character repaint test. Kept the background untouched. Supports a negative prompt, which let us name and block the mistakes it tended to make.Our default
FLUX.2 Dev (edit)About 8 of 9 once we matched the guidance setting. No negative prompt field, so every unwanted detail has to be fought in the main prompt.Good, harder to control
Qwen-Image-Edit 20BKept scenes perfectly, but wrote the made-up trigger word into the picture as lettering (10 of 25 images at 2,000 steps, 1 of 25 at 3,000). Colors drifted darker. Training cost 1.9× more.Only if you need it

The trap we almost fell into: our first Flux.2 Dev run looked much worse (4 of 9 clean) only because it ran at guidance 2.5 while Klein ran at 5. Re-running at 5 fixed four of the five failures. Never compare two LoRAs at different settings.

How many images, and what captions?

How many steps? Train long, save often, test every checkpoint

We trained to 5,000 steps, saved a checkpoint every 500, and scored each one from 1,500 on 43 test images the model never saw. "Fully clean" means correct anatomy, signature details kept, no invented lettering and nothing outside the character changed.

010203040281.5k312k332.5k293k313.5k324k304.5k265k
Fully clean results out of 43 test images, per checkpoint. The peak (2,500 and 4,000) is in the middle; 5,000 is the worst.
010203040371.5k392k392.5k383k393.5k414k384.5k375k
Correct hands out of 43. Best at 4,000 steps, then it slips.

What this means: the final checkpoint is rarely the best. On our style LoRA the same held: 4,000 and the final save beat 3,000, which was more scenic but dropped details. Budget time to test four to six checkpoints, starting around 1,500.

Why do the training previews look blurry?

Klein is a distilled model: the trainer draws its previews in 4 steps, so they look soft and smudged. Real use at 28 to 40 steps, guidance 4 to 5 is sharp. Don't throw away a LoRA because of its previews. Better still, switch previews off (they are slow on a 9B model) and test the saved checkpoints properly.

Settings that work for FLUX.2 Klein 9B

SettingValueWhy
TrainerOstris AI-ToolkitLoads Klein and Qwen; resumes cleanly after a crash
Model architectureflux2_klein_9bWithout it the trainer silently falls back to the wrong pipeline
Timestep typeshift (or sigmoid)weighted errors out on Klein
Precisionbf16, transformer not quantizedAbout 1.9 s per step, versus 5 to 9 s quantized
Save every500 steps, keep allYou'll want to test several, and resume if the job dies
Steps5,000Best checkpoints landed between 2,500 and 4,500

What hardware, and what does it cost?

How should you judge the results?

When the LoRA isn't the problem

Twice we blamed a LoRA and the fault was elsewhere. Characters kept losing the objects they held; the prompt that described them had been told to leave objects out. Once the prompt named the objects, they came back in 7 of 7 test images. And when the area the edit was allowed to repaint covered a prop, the model painted over the prop. Fix the mask and the prompt before you retrain.

The checklist

  1. Start on FLUX.2 Klein 9B.
  2. 20 to 45 clean, varied images. No duplicates or flips.
  3. Captions: trigger first, plain content, no style words, name the signature details.
  4. Train to 5,000 steps on an 80 GB GPU, save every 500.
  5. Ignore the previews. Test checkpoints from 1,500 up at 28 to 40 steps, guidance 4 to 5.
  6. Compare everything at identical settings, on a fixed test set.
  7. Check your judge against a human, and count what broke as well as what improved.

Want this set up for your studio's art? Book a call and bring a sample of your style.