MiniCPM5-2B-Math: a 2.5B model for math and proofs, with a proof swarm you can run yourself

English | 中文

MiniCPM5-2B-Math-GGUF

GGUF builds of MiniCPM5-2B-Math, a 2.5B math and proof specialist, for llama.cpp, Ollama and LM Studio. Pair it with the MiniCPM-Math Harness to run a small proof swarm on a laptop.

File Size Notes
MiniCPM5-2B-Math-Q4_K_M.gguf 1.6 GB smallest; quality loss likely on long proofs
MiniCPM5-2B-Math-Q8_0.gguf 2.7 GB recommended: usually close to BF16, and used for our laptop tests
MiniCPM5-2B-Math-BF16.gguf 5.0 GB full precision

Converted with llama.cpp b10964 (convert_hf_to_gguf.py, tokenizer pre-type minicpm5) and quantized with llama-quantize. The recommended sampling settings (temperature 0.9, top-p 0.95, min-p 0, top-k off) are embedded in the GGUF metadata, so llama.cpp uses them by default. All results on the model card were measured with the BF16 safetensors on vLLM.

File SHA-256
MiniCPM5-2B-Math-Q4_K_M.gguf 9ff4f86fd8eb3a16b1ff6b29d86366f3c963320e44a7325e6fb90c0142a7cf0b
MiniCPM5-2B-Math-Q8_0.gguf db251717a4d11b3221e7e2f1d096c3aebec94b1f25e8e24b4ebc56c185d32b9d
MiniCPM5-2B-Math-BF16.gguf 0703334ee732e110ef18dee427a8736fae52fbdd8a6ce3510de0df45dcaac72c

llama.cpp

llama-server -hf Caldalis/MiniCPM5-2B-Math-GGUF:Q8_0 --jinja -ngl 99 -c 98304 -np 4 -kvu --port 8000
  • --jinja uses the embedded chat template (thinking mode, <think> … </think>).
  • -np 4 -kvu gives 4 parallel slots that share a 96K-token KV cache: one slot per agent for the harness. Run the harness with --profile quick against this server.
  • --port 8000 is where the harness looks by default; without it, llama-server listens on 8080.
  • The embedded defaults already set min_p=0; llama.cpp's usual 0.05 can encourage repetition with this model.

Ollama

ollama run hf.co/Caldalis/MiniCPM5-2B-Math-GGUF:Q8_0

Start the server with a larger context window (for example OLLAMA_CONTEXT_LENGTH=65536 ollama serve), because proofs routinely need tens of thousands of tokens. With the harness, pass --base-url http://127.0.0.1:11434/v1 --model hf.co/Caldalis/MiniCPM5-2B-Math-GGUF:Q8_0 --no-raw.

Recommended settings

temperature=0.9, top_p=0.95, min_p=0.0, thinking enabled, and a large max_tokens. The published benchmark numbers let every call use the full 131,072-token context: a quarter of HMMT samples think past 64K tokens, and many of those still end with the right answer. On a laptop, 16K–32K is a practical compromise.

License

Apache-2.0 (see LICENSE).

Downloads last month
7
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Caldalis/MiniCPM5-2B-Math-GGUF

Quantized
(1)
this model