JSQ โ€” Qwen3-8B-Base (70% unstructured sparsity + INT4, symmetric)

Baseline compressed checkpoint for compression research.

  • Method: JSQ (joint sparsification + quantization, Wanda-style pruning + weight quant)
  • Sparsity: 70% unstructured
  • Quantization: INT4, per-group 128, symmetric (absmax)
  • Base: Qwen/Qwen3-8B-Base

Files

  • model.safetensors: sparse + fake-quantized weights in fp16, loadable via AutoModelForCausalLM.
  • compression/scales.safetensors: per-group-128 symmetric scales [out, in/128] per layer (scale=absmax/7).
  • compression_config.json: method / sparsity / bits / granularity / symmetric.

Notes

  • Mask recoverable as weight == 0. Symmetric: w=round(x/scale)*scale, levels in [-7,7].
Downloads last month
16
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support