Instructions to use contextboxai/Can-1.0-31B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use contextboxai/Can-1.0-31B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="contextboxai/Can-1.0-31B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("contextboxai/Can-1.0-31B") model = AutoModelForMultimodalLM.from_pretrained("contextboxai/Can-1.0-31B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use contextboxai/Can-1.0-31B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "contextboxai/Can-1.0-31B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "contextboxai/Can-1.0-31B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/contextboxai/Can-1.0-31B
- SGLang
How to use contextboxai/Can-1.0-31B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "contextboxai/Can-1.0-31B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "contextboxai/Can-1.0-31B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "contextboxai/Can-1.0-31B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "contextboxai/Can-1.0-31B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use contextboxai/Can-1.0-31B with Docker Model Runner:
docker model run hf.co/contextboxai/Can-1.0-31B
Cân 1.0 (31B)
Cân (Vietnamese: to weigh; cán cân = the scales of justice) is a decision model. Given a state (any text, JSON or conversation) and one typed question, it returns a calibrated probability distribution over the supplied options in one forward pass. Nothing is generated.
Question types follow TypeSafe /v1/systemone:
choice: pick one label fromcriteria→choice,probabilitiesscore: ordered scale (criteriais a list of levels) → expected levelscore,probabilitiesnoul: a statement that is true or false →noul= P(true)
| Base model | google/gemma-4-31B-it (Apache-2.0) |
| Fine-tune | LoRA r16 (merged), letter readout at the answer position |
| Training data | public human-labelled typed decisions (tasksource), procedural rule-application tasks with computed labels, synthetic workflow documents, programmatically generated math/code items with execution-verified labels |
| Teacher signal | soft labels from perplexity-ai/pplx-decider-v1.1-27b (Apache-2.0) mixed 50/50 with gold labels |
| Calibration | per (question type, option count) temperature, fitted on held-out non-benchmark data (calib.json) |
| Context | 1x H100: states up to 10k tokens read in full, longer ones truncated head+tail (never refused) |
Run
Tested configuration: 1x H100 80GB, vLLM 0.28.0.
pip install vllm==0.28.0 transformers
vllm serve contextboxai/Can-1.0-31B --served-model-name jevbeat --port 8890 \
--max-model-len 12288 --gpu-memory-utilization 0.95 --max-logprobs 20
# decision server (POST :8011/v1/systemone); states over JB_MAX_STATE tokens are truncated head 60% / tail 40%, never refused
JB_TOKENIZER=contextboxai/Can-1.0-31B JB_CALIB=calib.json JB_MAX_STATE=10000 python serve/server.py
Optional (not tested by us): with more GPU memory (e.g. --tensor-parallel-size 2 --max-model-len 98304), raise
JB_MAX_STATE to ~90000 to read the longest states in full.
Request: {"state": ..., "questions": {"decision": {"type": ..., "instructions": ..., "criteria": ...}}}.
Response: {"model", "answers": {"decision": {...}}, "usage": {"prompt_tokens", "completion_tokens", "total_tokens"}}.
More than 10 options -> HTTP 422.
Evaluation (internal proxy, not official)
Proxy slices (ours; not the official benchmark). Capability = (I + C) / 2 after calibration; paired doc-bootstrap where noted.
| slice | items | I | C | Capability | Quyet-1.0-Large (same harness, its own runtime) |
|---|---|---|---|---|---|
| TypeSafe evalsafe documents | 3030 | 57.3 | 90.9 | 74.1 | 68.8 |
| law / rules / math (LegalBench, MMLU-Pro, GSM8K) | 2295 | 74.4 | 90.1 | 82.2 | 84.7 |
| code (CRUXEval, HumanEval, MBPP; execution-verified) | 410 | 83.6 | 93.3 | 88.5 | 88.8 |
| multilingual (Global-MMLU, Belebele; 10 languages) | 500 | 68.5 | 89.2 | 78.9 | 81.6 |
| hand-written hard rules (fictional statutes; held out) | 100 | 68.4 | 69.6 | 69.0 | 70.1 |
Single forward pass, raw p50 latency 0.03 s on one H100 (vLLM).
Never trained on JevBench items or anything taken from the benchmark. Evaluation-only sources (LegalBench, MMLU-Pro, GSM8K, CRUXEval, HumanEval, MBPP, Global-MMLU, Belebele, TypeSafe evalsafe) were excluded from training and checked by 13-gram overlap.
Licence and credits
Apache-2.0. See NOTICE: Gemma 4 by Google (Apache-2.0); pplx-decider-v1.1-27b by Perplexity (Apache-2.0, used as a
teacher); tasksource by Damien Sileo (per-dataset licences; non-commercial rows excluded).
Built by ContextBox AI.
- Downloads last month
- 51