Cân 1.0 (31B)

Cân (Vietnamese: to weigh; cán cân = the scales of justice) is a decision model. Given a state (any text, JSON or conversation) and one typed question, it returns a calibrated probability distribution over the supplied options in one forward pass. Nothing is generated.

Question types follow TypeSafe /v1/systemone:

  • choice: pick one label from criteria → choice, probabilities
  • score: ordered scale (criteria is a list of levels) → expected level score, probabilities
  • noul: a statement that is true or false → noul = P(true)
Base model google/gemma-4-31B-it (Apache-2.0)
Fine-tune LoRA r16 (merged), letter readout at the answer position
Training data public human-labelled typed decisions (tasksource), procedural rule-application tasks with computed labels, synthetic workflow documents, programmatically generated math/code items with execution-verified labels
Teacher signal soft labels from perplexity-ai/pplx-decider-v1.1-27b (Apache-2.0) mixed 50/50 with gold labels
Calibration per (question type, option count) temperature, fitted on held-out non-benchmark data (calib.json)
Context 1x H100: states up to 10k tokens read in full, longer ones truncated head+tail (never refused)

Run

Tested configuration: 1x H100 80GB, vLLM 0.28.0.

pip install vllm==0.28.0 transformers
vllm serve contextboxai/Can-1.0-31B --served-model-name jevbeat --port 8890 \
  --max-model-len 12288 --gpu-memory-utilization 0.95 --max-logprobs 20
# decision server (POST :8011/v1/systemone); states over JB_MAX_STATE tokens are truncated head 60% / tail 40%, never refused
JB_TOKENIZER=contextboxai/Can-1.0-31B JB_CALIB=calib.json JB_MAX_STATE=10000 python serve/server.py

Optional (not tested by us): with more GPU memory (e.g. --tensor-parallel-size 2 --max-model-len 98304), raise JB_MAX_STATE to ~90000 to read the longest states in full.

Request: {"state": ..., "questions": {"decision": {"type": ..., "instructions": ..., "criteria": ...}}}. Response: {"model", "answers": {"decision": {...}}, "usage": {"prompt_tokens", "completion_tokens", "total_tokens"}}. More than 10 options -> HTTP 422.

Evaluation (internal proxy, not official)

Proxy slices (ours; not the official benchmark). Capability = (I + C) / 2 after calibration; paired doc-bootstrap where noted.

slice items I C Capability Quyet-1.0-Large (same harness, its own runtime)
TypeSafe evalsafe documents 3030 57.3 90.9 74.1 68.8
law / rules / math (LegalBench, MMLU-Pro, GSM8K) 2295 74.4 90.1 82.2 84.7
code (CRUXEval, HumanEval, MBPP; execution-verified) 410 83.6 93.3 88.5 88.8
multilingual (Global-MMLU, Belebele; 10 languages) 500 68.5 89.2 78.9 81.6
hand-written hard rules (fictional statutes; held out) 100 68.4 69.6 69.0 70.1

Single forward pass, raw p50 latency 0.03 s on one H100 (vLLM).

Never trained on JevBench items or anything taken from the benchmark. Evaluation-only sources (LegalBench, MMLU-Pro, GSM8K, CRUXEval, HumanEval, MBPP, Global-MMLU, Belebele, TypeSafe evalsafe) were excluded from training and checked by 13-gram overlap.

Licence and credits

Apache-2.0. See NOTICE: Gemma 4 by Google (Apache-2.0); pplx-decider-v1.1-27b by Perplexity (Apache-2.0, used as a teacher); tasksource by Damien Sileo (per-dataset licences; non-commercial rows excluded).

Built by ContextBox AI.

Downloads last month
51
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for contextboxai/Can-1.0-31B

Finetuned
(292)
this model
Quantizations
1 model