W2-31B-A3B-dLLM-Base
Highlights
- 31.1B total parameters with a 3.2B active parameter path
- Generates 8–32 tokens per step before self-distillation.
- The model natively self-edits through iterative denoising.
- Sparse MoE design with 8 routed experts and 1 shared expert active per token
- Grouped-query attention with QK normalization
- BF16 and routed-expert W4A16 weight variants
Model Overview
- Type: Native Diffusion Transformer (DiT) large language model
- Architecture: Sparse Mixture-of-Experts (MoE)
- Total Parameters: 31.1B
- Activated Parameters: 3.2B
- Hidden Dimension: 2,048
- Layers: 48
- Attention: GQA with 32 query heads and 4 key-value heads
- Experts: 8 activated out of 144 routed experts, plus 1 shared expert
- Position Embedding: RoPE
- Embeddings and Output Head: Untied
- Vocabulary Size: 129,024
Model Weights
- BF16 weights: repository root
- INT4 routed-expert weights (W4A16):
int4/
- Downloads last month
- 7