You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

W2-31B-A3B-dLLM-Base

WhaleTech banner
W2 dLLM Demo
W2 dLLM Demo (English)

Highlights

  • 31.1B total parameters with a 3.2B active parameter path
  • Generates 8–32 tokens per step before self-distillation.
  • The model natively self-edits through iterative denoising.
  • Sparse MoE design with 8 routed experts and 1 shared expert active per token
  • Grouped-query attention with QK normalization
  • BF16 and routed-expert W4A16 weight variants

Model Overview

  • Type: Native Diffusion Transformer (DiT) large language model
  • Architecture: Sparse Mixture-of-Experts (MoE)
  • Total Parameters: 31.1B
  • Activated Parameters: 3.2B
  • Hidden Dimension: 2,048
  • Layers: 48
  • Attention: GQA with 32 query heads and 4 key-value heads
  • Experts: 8 activated out of 144 routed experts, plus 1 shared expert
  • Position Embedding: RoPE
  • Embeddings and Output Head: Untied
  • Vocabulary Size: 129,024

Model Weights

  • BF16 weights: repository root
  • INT4 routed-expert weights (W4A16): int4/
Downloads last month
7
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support