JSQ โ Qwen3-8B-Base (70% unstructured sparsity + INT4, symmetric)
Baseline compressed checkpoint for compression research.
- Method: JSQ (joint sparsification + quantization, Wanda-style pruning + weight quant)
- Sparsity: 70% unstructured
- Quantization: INT4, per-group 128, symmetric (absmax)
- Base: Qwen/Qwen3-8B-Base
Files
model.safetensors: sparse + fake-quantized weights in fp16, loadable viaAutoModelForCausalLM.compression/scales.safetensors: per-group-128 symmetric scales[out, in/128]per layer (scale=absmax/7).compression_config.json: method / sparsity / bits / granularity / symmetric.
Notes
- Mask recoverable as
weight == 0. Symmetric:w=round(x/scale)*scale, levels in[-7,7].
- Downloads last month
- 16
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support