Training LoRAs on FLUX.2 Klein: What We Learned in 2026
Short answer: for game art in 2026, train your LoRA on FLUX.2 Klein 9B, use 20 to 45 clean images with short captions, train to about 5,000 steps, save every 500, and pick the checkpoint by testing, not by the training previews. In our runs the best checkpoint was usually around 4,000 steps, and the last one was never the best.
Two years ago we shared what 50 Flux Dev models taught us. Since then we moved our production training to FLUX.2 Klein and trained more than a dozen LoRAs for game studios: mascots, character sets, UI panels, reward icons and a before-and-after "edit" LoRA that repaints a character in place. This is what we learned, with the numbers.
Which base model should you train on?
We trained the same jobs on three bases and compared them on identical test images, prompts and seeds.
| Base model | What happened | Verdict |
|---|---|---|
| FLUX.2 Klein 9B | 9 of 9 clean on a character repaint test. Kept the background untouched. Supports a negative prompt, which let us name and block the mistakes it tended to make. | Our default |
| FLUX.2 Dev (edit) | About 8 of 9 once we matched the guidance setting. No negative prompt field, so every unwanted detail has to be fought in the main prompt. | Good, harder to control |
| Qwen-Image-Edit 20B | Kept scenes perfectly, but wrote the made-up trigger word into the picture as lettering (10 of 25 images at 2,000 steps, 1 of 25 at 3,000). Colors drifted darker. Training cost 1.9× more. | Only if you need it |
The trap we almost fell into: our first Flux.2 Dev run looked much worse (4 of 9 clean) only because it ran at guidance 2.5 while Klein ran at 5. Re-running at 5 fixed four of the five failures. Never compare two LoRAs at different settings.
How many images, and what captions?
- 20 to 45 images is enough. Our strongest style LoRA used 45; our character LoRAs 40 to 44. More steps, not more images, is what makes the model remember.
- Don't pad the set with flips, crops or near-duplicates. The same image twice teaches the model to copy it.
- Caption format: trigger word first, then a plain description of what is in the picture: "xopax, a treasure chest". No style words like "glossy" or "cartoon". The style is what the LoRA learns, so describing it steals it from the trigger.
- Name the small signature details in every caption ("a lighter spiral on the chest"). When we left a chest mark unnamed, the model started turning it into letters.
- Pick a trigger word that doesn't look like a word. Letter-like tokens such as xpscrx sometimes get written into the image as text. Test for it, and if it happens, switch to a descriptive phrase.
How many steps? Train long, save often, test every checkpoint
We trained to 5,000 steps, saved a checkpoint every 500, and scored each one from 1,500 on 43 test images the model never saw. "Fully clean" means correct anatomy, signature details kept, no invented lettering and nothing outside the character changed.
What this means: the final checkpoint is rarely the best. On our style LoRA the same held: 4,000 and the final save beat 3,000, which was more scenic but dropped details. Budget time to test four to six checkpoints, starting around 1,500.
Why do the training previews look blurry?
Klein is a distilled model: the trainer draws its previews in 4 steps, so they look soft and smudged. Real use at 28 to 40 steps, guidance 4 to 5 is sharp. Don't throw away a LoRA because of its previews. Better still, switch previews off (they are slow on a 9B model) and test the saved checkpoints properly.
Settings that work for FLUX.2 Klein 9B
| Setting | Value | Why |
|---|---|---|
| Trainer | Ostris AI-Toolkit | Loads Klein and Qwen; resumes cleanly after a crash |
| Model architecture | flux2_klein_9b | Without it the trainer silently falls back to the wrong pipeline |
| Timestep type | shift (or sigmoid) | weighted errors out on Klein |
| Precision | bf16, transformer not quantized | About 1.9 s per step, versus 5 to 9 s quantized |
| Save every | 500 steps, keep all | You'll want to test several, and resume if the job dies |
| Steps | 5,000 | Best checkpoints landed between 2,500 and 4,500 |
What hardware, and what does it cost?
- Use an 80 GB GPU (A100 80GB) for Klein 9B. Loading the model and its text encoder needs over 50 GB of system memory; 48 GB machines get killed during load with no error message.
- A normal character or style LoRA: about 1.9 s per step, so 5,000 steps take about 2.6 hours and cost roughly $4 to $6 on a rented cloud GPU.
- An edit LoRA (trained on before-and-after pairs) is about 2.4× slower, because every example carries two images. Our 5,000-step edit run took 7.7 hours and cost about $9.
- Protect long runs. Our cloud host killed the training process twice in one night from outside the machine. Run it under a small wrapper that ignores that signal, and rely on the trainer's auto-resume from the last saved checkpoint. Copy each checkpoint off the machine as it lands, so a stalled disk costs you steps, not the run.
How should you judge the results?
- Use a fixed test set the model never saw, the same prompt and seed for every checkpoint, and look at the same images side by side.
- If you use an AI judge, check it against a human first. Our first automatic "hands" checker counted fingers; the art director was judging whether the hand read as a soft mitten. They agreed only 61% of the time. We rewrote the checker around what the director actually looks at.
- Check what the LoRA broke, not just what it fixed: count the images where anything outside the subject changed. Our edit LoRA left everything outside the character untouched on all 43 test images; the LoRA we had in production changed something outside it on 22 of 43.
When the LoRA isn't the problem
Twice we blamed a LoRA and the fault was elsewhere. Characters kept losing the objects they held; the prompt that described them had been told to leave objects out. Once the prompt named the objects, they came back in 7 of 7 test images. And when the area the edit was allowed to repaint covered a prop, the model painted over the prop. Fix the mask and the prompt before you retrain.
The checklist
- Start on FLUX.2 Klein 9B.
- 20 to 45 clean, varied images. No duplicates or flips.
- Captions: trigger first, plain content, no style words, name the signature details.
- Train to 5,000 steps on an 80 GB GPU, save every 500.
- Ignore the previews. Test checkpoints from 1,500 up at 28 to 40 steps, guidance 4 to 5.
- Compare everything at identical settings, on a fixed test set.
- Check your judge against a human, and count what broke as well as what improved.
Want this set up for your studio's art? Book a call and bring a sample of your style.