Krea 2 Turbo β€” 4-Step Distillation LoRA

Half the steps Β· ~1.6Γ— faster Β· 45% of the gap to the 8-step teacher closed Β· texture at 1.03Γ— the teacher's, verified clean Β· teacher preferred on only 6 of 45 judged renders Β· 12 trained resolutions Β· 78,000 training samples Β· 21 days on one RTX 3090.

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4 β€” Turbo's own model and sigmas, half the denoising passes, and the fine texture that 4-step Turbo loses put back.

  • ⚑ Half the steps β€” 8 β†’ 4, on Turbo's own deployment sigmas
  • ⏱️ ~1.6Γ— faster end to end β€” 54.5 s vs 88.7 s at 1024Γ—1024 (1.8Γ— on denoise alone)
  • 🎯 Texture as good as the teacher or better β€” 1.03Γ— the teacher's fine-texture energy at 1280Β² and 1440Β², every frequency band within 10%; verified clean: saturation 0.96–0.97Γ—, fewer clipped highlights/shadows, skin texture 0.97–0.99Γ—
  • πŸ“ 45% of the 4-step gap closed β€” held-out velocity error fell from 4.70e-02 (stock Turbo at 4 steps) to 2.59e-02 with the LoRA; VLM judge preferred teacher on only 6 of 45 renders (37 ties, 2 LoRA wins)
  • πŸ—£οΈ Prompt-aware training β€” critic scores images against their prompts during training
  • πŸ“ 12 trained resolutions β€” multi-aspect from 512Γ—512 up to 1440Γ—1440
  • πŸ”Œ Drop-in β€” plain LoRA weights for diffusers and ComfyUI. No custom nodes, no patched sampler, no code
  • 🎲 13,750 prompts from Lakonik's 3M-prompt dataset, each recorded as a full 8-step teacher trajectory
  • πŸ“· 43,044 real-photo crops β€” 25,560 from LSDIR + 17,484 from Flickr2K, captioned, in critic's real set
  • πŸ”’ 78,000 training samples in shipped weights
  • πŸ“… 21 days on a single RTX 3090
  • πŸ” 16 recipe adjustments β€” each made on measurement of the one before

πŸ’‘ Too strong on a prompt? Turn it down. The LoRA restores fine texture, and on some subjects β€” stylised art, high-contrast splash pieces, very large renders β€” that can read as too much at strength 1.0. Effect scales smoothly: 0.75 is a good second setting, 0.6–0.9 is fair game.

The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 4 steps


Files

file what it is
krea2_turbo_4step_rank_64_lora.safetensors LoRA in diffusers key format β€” see diffusers
krea2_turbo_4step_rank_64_lora_comfyui.safetensors Same weights under ComfyUI key names β€” see ComfyUI
krea2_turbo_4step_lora_t2i.json Ready ComfyUI workflow, stock nodes only
LICENSE.pdf Krea 2 Community License Agreement
NOTICE.txt Required attribution notice

Both weight files are one adapter β€” only key names differ. Both carry training details in safetensors metadata.


Quick Start

diffusers

pip install git+https://github.com/huggingface/diffusers.git
import torch
from diffusers import Krea2Pipeline
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

pipe = Krea2Pipeline.from_pretrained("krea/Krea-2-Turbo", torch_dtype=torch.bfloat16).to("cuda")

lora = hf_hub_download("lvladikov/Krea2-Turbo-Distill-4step-LoRA", "krea2_turbo_4step_rank_64_lora.safetensors")
state = load_file(lora)
state = {f"transformer.{k}": v for k, v in state.items() if not k.endswith(".alpha")}
pipe.load_lora_weights(state, adapter_name="4step")

image = pipe("a fox in the snow", num_inference_steps=4, guidance_scale=0.0).images[0]
image.save("krea2_4step.png")
  • num_inference_steps=4 is the whole config. Pipeline applies Turbo's fixed timestep shift (ΞΌ = 1.15) and evaluates at Οƒ = 1.0, 0.905, 0.760, 0.513 β€” exactly the four points the LoRA was trained on. Keep guidance_scale=0.0.
  • Strength: pipe.set_adapters(["4step"], adapter_weights=[0.75]). Stock Turbo for comparison: pipe.unload_lora_weights() + num_inference_steps=8.

ComfyUI

file put it in
krea2_turbo_4step_rank_64_lora_comfyui.safetensors ComfyUI/models/loras/
krea2_turbo_bf16.safetensors (from Comfy-Org/Krea-2) ComfyUI/models/diffusion_models/
qwen3vl_4b_bf16.safetensors (same repo) ComfyUI/models/text_encoders/
qwen_image_vae.safetensors (same repo) ComfyUI/models/vae/

Load krea2_turbo_4step_lora_t2i.json. Full bf16, no quantisation, runs on CUDA / Apple Silicon / CPU unchanged.

Settings: steps 4, cfg 1.0, sampler euler / simple, LoRA strength 1.0.

βš™οΈ cfg 1.0, not 0.0. ComfyUI expresses "no CFG" as 1.0 (one forward pass); diffusers uses 0.0. Setting 0.0 in ComfyUI is not the same thing.


Performance (1024Γ—1024, Apple Silicon MLX bf16)

load denoise total
Turbo 8 steps (quality bar) 8.2 s 77.5 s 88.7 s
Turbo 4 steps, no LoRA 7.8 s 38.8 s 49.7 s
Turbo 4 steps + this LoRA 7.3 s 44.0 s 54.5 s

4 steps with LoRA is ~1.6Γ— faster than the 8-step bar (54.5 s vs 88.7 s). Denoise alone: 1.8Γ— (44.0 s vs 77.5 s).

Per-step cost: 9.7 s β†’ 11.0 s (~13% slower), plus ~1.3 GB peak memory. Loading the LoRA costs nothing measurable. Halving steps wins comfortably.


LoRA Strength

strength what happens
below 1.0 Correction partly applied β€” output between unassisted 4-step and full LoRA. 0.75 for when 1.0 is too textured/hard
1.0 Trained point, recommended
1.0–1.5 Extrapolation β€” surface detail denser, micro-contrast harder, coherent but stylised
above 1.5 Not recommended β€” breaks into uniform speckle/noise over whole image

Reach for steps before strength: 1.0 at more steps is the dependable way to get more. Strength and step count trade against each other.

LoRA strength comparison


Method

Progressive distillation with Krea 2 Turbo as its own teacher. Teacher ran 8 steps at ΞΌ = 1.15, guidance 0.0; full trajectory recorded. Student trained to cover two teacher steps in one:

v_target = (x_{i+2} βˆ’ x_i) / (Οƒ_{i+2} βˆ’ Οƒ_i)

The even indices of the 8-step schedule are precisely the four sigmas the 4-step student deploys on β€” no interpolation, no schedule mismatch.

Critic (LADD-style) on frozen Turbo features at block 14:

  • Real: teacher finals (unpaired) or real photographs (half the draws), re-noised to Οƒ ∈ [0.02, 0.5]
  • Fake: student's predicted clean latent
  • Prompt-aware head reads pooled text vector; mismatch term forces real images under wrong prompts to read fake
  • Real photos captioned so they participate in prompt-aware term

Three key weight choices:

  • Final chord (Οƒ = 0.512, decides fine texture) weighted 3Γ— in PD loss
  • Low-frequency anchor on large buckets (weight 0.2) β€” structure held to teacher, fine band free
  • Shipped adapter is Polyak (EMA) average (decay 0.999), not last live state

What the LoRA Touches

Rank 64, alpha = rank (scale 1.0), bf16. 228 modules:

  • 224 block linears β€” all 28 transformer blocks: to_q, to_k, to_v, to_gate, to_out.0, ff.gate, ff.up, ff.down
  • 4 global linears β€” time_embed.linear_1, time_embed.linear_2, time_mod_proj, final_layer.linear

The global linears are included deliberately. Measuring Krea's own Raw→Turbo delta:

layer relative β€–Ξ”Wβ€–/β€–Wβ€–
time_embed.linear_2 0.0777 ← largest change in network
time_embed.linear_1 0.0429
final_layer.linear 0.0265
typical block linear ~0.014

time_embed.linear_2 moves ~5.5Γ— more than any block linear. Changing step count is largely a change to how the model reads the timestep β€” a LoRA freezing the timestep path withholds the weights the task most needs.


Training Data

  • 13,750 prompts from Lakonik/t2i-prompts-3m (sampled without replacement, deduplicated, filtered), one recorded teacher trajectory each
  • 203 held-out validation + 641 OOD evaluation β€” never received a gradient step
  • 43,044 real-photo crops β€” 25,560 LSDIR + 17,484 Flickr2K, cut at native resolution to 12 buckets, quality-gated, VAE-encoded, captioned per crop

Photos only reached the critic β€” never regression targets β€” so adapter learned detail density from reality, not content.


Resolutions (12 buckets)

512Γ—512 512Γ—768 768Γ—512
768Γ—768 768Γ—1024 1024Γ—768
1024Γ—1024 960Γ—1280 1280Γ—960
1280Γ—1280 1440Γ—1280 1440Γ—1440

1440Γ—1440 had a deliberately small share β€” enough to learn the size without spending budget there. Buckets interleaved by remaining samples, not curriculum.


Hardware

Trained on single RTX 3090 (24 GB). Frozen base quantized weight-only to int8 (blockwise-64) β€” ~0.007 relative error, ~3% step time cost. Teacher trajectories rolled at same precision (targets permanent, error baked in).

Every bucket trained under full int8 up to 1440Γ—1440. Large buckets fit via:

  • Checkpointed block inputs staged to pinned host memory (numerically exact)
  • Discriminator pass after generator backward (peaks don't overlap)
  • Trunk stopped at feature tap it actually reads

Released LoRA is bf16 applied to unquantized base.


Usage Notes

  • 🎯 Krea 2 Turbo only β€” trained against Turbo's weights and schedule
  • 🚫 Keep guidance at 0.0 (ComfyUI: cfg 1.0, not 0.0)
  • πŸ“ Keep ΞΌ = 1.15 β€” training targets anchored to that grid
  • πŸ”¬ Training used int8 base; released LoRA is bf16 on unquantized base

Examples

Every sheet below: base model (8 steps), base at 4 steps without LoRA, base at 4 steps with LoRA β€” same seed throughout. Compare panels 2 vs 3 to isolate LoRA effect. NFE = steps (Turbo is CFG-free).

Portrait of a young woman with freckles and windswept auburn hair...

portrait comparison

Turbo 8 steps Turbo 4 steps, no LoRA Turbo 4 steps + LoRA
8 steps 4 steps 4 steps + LoRA

Kingfisher bursting out of water...

kingfisher comparison

Turbo 8 steps Turbo 4 steps, no LoRA Turbo 4 steps + LoRA
8 steps 4 steps 4 steps + LoRA

Rainy night city street with glowing neon signs...

neonstreet comparison

Turbo 8 steps Turbo 4 steps, no LoRA Turbo 4 steps + LoRA
8 steps 4 steps 4 steps + LoRA

(12 more examples in full version β€” see assets/ for all 15 prompts)


Resolution Sweeps

assets/resolution_sweeps/ β€” this LoRA at every trained resolution for all 15 test prompts (same prompts, seed, 4 steps, strength 1.0). Nothing cherry-picked.

_teacher-8step/ β€” official Krea 2 Turbo 8-step reference renders for same prompts/seeds/resolutions.

Layout:

assets/resolution_sweeps/
β”œβ”€β”€ _teacher-8step/         8-step stock Turbo reference
β”‚   β”œβ”€β”€ 512x512/
β”‚   β”œβ”€β”€ 1024x1024/
β”‚   └── ... (all 12 buckets)
β”œβ”€β”€ 4step-LoRA/             Same tree, this LoRA at 4 steps
└── 2step-LoRA-extreme/     Out-of-spec 2-step strips (bonus section)

Two ways to read: down the sweep (does it hold across resolutions?) or against teacher (open same file in both folders).


Bonus: 2-Step Extreme Test

Out-of-spec experiment β€” not for production. Useful as fast preview: at 2 steps, LoRA at strength 1.0 gives reliable read on composition/look at quarter of 8 steps.

Stock Turbo ghosts/smears at 2 steps; with LoRA image stays coherent and sharp. All 15 prompts checked: extra detail is real subject detail (hair, skin, fabric) + sharper atmospheric elements where prompt calls for it. Nothing unprompted appears, nothing intrudes on faces.

Full strips: assets/resolution_sweeps/2step-LoRA-extreme/

πŸ”­ Dedicated 2-step LoRA now published: lvladikov/Krea2-Turbo-Distill-2step-LoRA. Renders at 2 steps for fast previews/drafts. Not this adapter's quality β€” for quality renders, stay here.


Archive

Earlier checkpoints and sweeps under _archive/ β€” superseded, not maintained.


What's next

A 2-step LoRA was the natural follow-on, and it is now published as its own project: lvladikov/Krea2-Turbo-Distill-2step-LoRA. It is a separate release rather than a competitor to this one: at two steps a distilled model gives up more than at four, so the aim is a usable 2-step Turbo β€” fast previews and drafts at half this adapter's cost, a quarter of the teacher's β€” not the quality bar this adapter holds. It is still in training, and its page carries the measurements and the faults as they stand.


Detailed Model Card

For more details, if interested, have a look at the Detailed Model Card.


License

This adapter is a Derivative of Krea 2 Turbo under the Krea 2 Community License Agreement. Everything the agreement says about Krea 2 Turbo applies to this LoRA: Acceptable Use Policy, revenue threshold for commercial use, content-filtering duty for deployments.

Copy of agreement: LICENSE.pdf, required notice: NOTICE.txt. See krea.ai/krea-2-licensing.

Not an official Krea product, not endorsed by Krea. Base model is Krea's; adapter weights and everything else in this repo are mine.

Downloads last month
27,471
Inference Providers NEW

Model tree for lvladikov/Krea2-Turbo-Distill-4step-LoRA

Base model

krea/Krea-2-Raw
Adapter
(1640)
this model
Adapters
2 models

Space using lvladikov/Krea2-Turbo-Distill-4step-LoRA 1