UNet

Flow matching (time_shift_type is 'linear') and the LFM2.5 text encoder on the top of Aniimage.

Due to the lack of training data, the only prompt it understands is anime portraits.

VAE

This release is an adaptation of the 8BitStudio's model for the Mage-VAE.

The UNet takes a 16x downsampled latent, which is smaller than SDXL thanks to the new autoencoder. The decoding is 5x faster than Flux.2 VAE.

There's the LFM2.5 encoder, which is great for those anime text sequences.

The learning rate was set to 1e-5 during the warmup phase, which had 200k images. Then it continued with a fixed image size until the handwritten code was confirmed to converge in later epochs.

For the second stage, the loss was calculated by the difference between the latent, the colored images and the grayscale/colored image pairs; in mixed-resolution image samples.

The timesteps were chosen by the logit-normal sampling.

Source data:

  • anime_faces_256px_v2
  • anime_style_portrait
  • gelbooru (landscape)
  • portraits_512
  • wikiart_face
Downloads last month
16
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nebulette/aniportrait-lfm

Finetuned
(1)
this model