dpsk-v4-flash.h070.sft4_step2600.step_2600

AgentPTB sweep checkpoint. Cell dpsk-v4-flash โ€” pi / DeepSeek v4-flash @ effort thinking.

field value
plot cell dpsk-v4-flash
driver pi / DeepSeek v4-flash
reasoning effort thinking
run boot (UTC) 2026-08-11T08:38:35Z
role intermediate
hours into run h70.65 of 100
checkpoint path in run ckpts/sft4_step2600/step_2600
shards 1
size 18.8 GB
base model Qwen/Qwen3.5-9B-Base
eos_token_id [248044] โš ๏ธ MISSING 248046

Reading the eos field

248046 is <|im_end|>, the token the Qwen3.5 chat template ends every assistant turn with. Checkpoints missing it do not stop at end-of-turn and overrun the context window, so their eval numbers are a floor, not a measurement โ€” compare them only against other checkpoints with the same eos status, or re-package before evaluating.

Mapping back to the figures

The repo id is {cell}.h{HHH}.{family}.{step}, where hHHH is the hour of the 100-hour run at which this checkpoint was written โ€” the same x-axis the sweep figures use for eval panels (t_h). So a checkpoint drops onto the performance-over-time curve directly, and sorting repo ids within a cell sorts them chronologically.

hHHH is rounded down to whole hours for sortability; the exact value is the hours into run row above, and in agentic-ptb/INDEX.

Downloads last month
8
Safetensors
Model size
9B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for agentic-ptb/dpsk-v4-flash.h046.sft3_final.step_1900

Finetuned
(584)
this model