Instructions to use soyrsoyr/Qwen3.6-27B-FP8-dynamic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use soyrsoyr/Qwen3.6-27B-FP8-dynamic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="soyrsoyr/Qwen3.6-27B-FP8-dynamic") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("soyrsoyr/Qwen3.6-27B-FP8-dynamic") model = AutoModelForMultimodalLM.from_pretrained("soyrsoyr/Qwen3.6-27B-FP8-dynamic", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use soyrsoyr/Qwen3.6-27B-FP8-dynamic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "soyrsoyr/Qwen3.6-27B-FP8-dynamic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "soyrsoyr/Qwen3.6-27B-FP8-dynamic", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/soyrsoyr/Qwen3.6-27B-FP8-dynamic
- SGLang
How to use soyrsoyr/Qwen3.6-27B-FP8-dynamic with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "soyrsoyr/Qwen3.6-27B-FP8-dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "soyrsoyr/Qwen3.6-27B-FP8-dynamic", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "soyrsoyr/Qwen3.6-27B-FP8-dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "soyrsoyr/Qwen3.6-27B-FP8-dynamic", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use soyrsoyr/Qwen3.6-27B-FP8-dynamic with Docker Model Runner:
docker model run hf.co/soyrsoyr/Qwen3.6-27B-FP8-dynamic
Qwen3.6-27B-FP8-dynamic
FP8-dynamic (W8A8) quantization of Qwen/Qwen3.6-27B, produced with llm-compressor and served with vLLM.
Model overview
- Base model: Qwen/Qwen3.6-27B (Qwen3.5-family hybrid-attention VLM)
- Quantization scheme:
FP8_DYNAMIC(weights quantized to FP8 (E4M3), per-channel, round-to-nearest; activations quantized to FP8 per-token with dynamic scales computed at runtime, no calibration data required). - What's quantized: the text transformer's
Linearlayers. The vision tower (re:.*visual.*), the hybrid linear-attention layers (re:.*linear_attn.*), the MoE router (re:.*mlp.gate), andlm_headstay in the original precision. - Format: compressed-tensors.
The full VLM (with
vision_config) is saved, so it loads directly in vLLM. - Memory: ~2x smaller than the bf16 base for the quantized weights.
Deployment (vLLM)
vllm serve soyrsoyr/Qwen3.6-27B-FP8-dynamic
from vllm import LLM, SamplingParams
llm = LLM(model="soyrsoyr/Qwen3.6-27B-FP8-dynamic")
out = llm.generate(["The capital of France is"], SamplingParams(max_tokens=32))
print(out[0].outputs[0].text)
Creation
Quantized with llm-compressor. Because current vLLM only registers the
multimodal Qwen3_5ForConditionalGeneration architecture, the model is loaded
with AutoModelForImageTextToText and the full model is saved (text
quantized, vision tower kept in bf16).
from transformers import AutoModelForImageTextToText, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
MODEL_ID = "Qwen/Qwen3.6-27B"
model = AutoModelForImageTextToText.from_pretrained(MODEL_ID, dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8_DYNAMIC",
ignore=["lm_head", "re:.*visual.*", "re:.*linear_attn.*", "re:.*mlp.gate$"],
)
oneshot(model=model, recipe=recipe) # data-free
model.save_pretrained("Qwen3.6-27B-FP8-dynamic", save_compressed=True, save_original_format=False)
tokenizer.save_pretrained("Qwen3.6-27B-FP8-dynamic")
Evaluation
Evaluated with lm-evaluation-harness (OpenLLM Leaderboard v1) and lighteval (generative reasoning suite), both via the vLLM backend.
OpenLLM Leaderboard v1
| Benchmark (OpenLLM v1) | Metric | Qwen3.6-27B (base) | FP8-dynamic | Recovery |
|---|---|---|---|---|
| ARC-Challenge (25-shot) | acc_norm | 48.04 | 49.06 | 102.1% |
| HellaSwag (10-shot) | acc_norm | 71.25 | 71.63 | 100.5% |
| TruthfulQA-mc2 (0-shot) | acc | 55.06 | 55.65 | 101.1% |
| Winogrande (5-shot) | acc | 77.19 | 77.43 | 100.3% |
| Average | 62.89 | 63.44 | 100.9% |
FP8-dynamic recovers 100.9% of the base model's average score on these tasks, i.e. quantization is effectively lossless here.
MMLU and GSM8k are omitted: Qwen3.6-27B is a reasoning ("thinking") model, and
under the OpenLLM v1 protocol both collapse to a harness artifact rather than a
real score. GSM8k is generative and its chain-of-thought is truncated by the
default max_gen_toks before the final answer, and MMLU's letter-answer
loglikelihood is measured at the position where the model wants to emit its
<think> block, so every subject lands at random chance (MMLU scored 23.02 for
the base model and 23.05 for FP8, both at the 25% four-way chance level, which
confirms it is a measurement artifact and not quantization damage). Math and
knowledge are better captured by the generative reasoning suite (AIME, MATH-500,
GPQA).
Reasoning suite (generative)
The eval that actually reflects a thinking model: single-sample pass@1, model allowed to reason (temperature 0.6, top_p 0.95, up to 32k generated tokens), scored with lighteval.
| Benchmark | Metric | Qwen3.6-27B (base) | FP8-dynamic | Recovery |
|---|---|---|---|---|
| AIME-24 | pass@1 | 83.33 | 73.33 | 88.0% |
| AIME-25 | pass@1 | 73.33 | 73.33 | 100.0% |
| MATH-500 | pass@1 | 86.60 | 86.20 | 99.5% |
| Average | 81.09 | 77.62 | 95.7% |
AIME-25 and MATH-500 are effectively lossless (100.0% and 99.5%). The AIME-24 number is a single-sample (pass@1, n=1) score over only 30 problems, so its standard error is large (~7 points, i.e. roughly a 2-problem swing); the 10-point gap there is within that sampling noise rather than clear quantization damage. For a lower-variance figure one would average over multiple samples (e.g. avg@64), which was out of scope for this single-GPU run. GPQA-diamond will be included once dataset access is granted (it is a gated dataset on the Hub).
Reproduction
OpenLLM v1 (lm-evaluation-harness):
lm_eval --model vllm \
--model_args pretrained=soyrsoyr/Qwen3.6-27B-FP8-dynamic,dtype=auto,gpu_memory_utilization=0.8,max_model_len=16384,max_num_seqs=256 \
--tasks arc_challenge --num_fewshot 25 --apply_chat_template --fewshot_as_multiturn --batch_size auto
# ...repeated for hellaswag (10), truthfulqa_mc2 (0), winogrande (5)
Reasoning suite (lighteval), scored pass@1 with the model's own chat template:
lighteval vllm \
"model_name=soyrsoyr/Qwen3.6-27B-FP8-dynamic,dtype=auto,max_model_length=40960,generation_parameters={temperature:0.6,top_p:0.95,top_k:20,max_new_tokens:32768}" \
"aime24|0,aime25|0,math_500|0" --save-details
- Downloads last month
- 19
Model tree for soyrsoyr/Qwen3.6-27B-FP8-dynamic
Base model
Qwen/Qwen3.6-27B