--- license: apache-2.0 base_model: BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT datasets: - BytedTsinghua-SIA/Open-MOPD-Data library_name: transformers pipeline_tag: text-generation language: - en tags: - smollm3 - open-mopd - reinforcement-learning - code --- # Open-MOPD-SmolLM3-3B-RL-Code This is the code-domain teacher in the Open-MOPD pipeline. It starts from `BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT` and is trained only on code prompts with verifiable rewards using GRPO. This release corresponds to training step 180. Training uses global batch size 128, mini-batch size 32, learning rate `1e-6`, rollout group size 16, a 30,000-token response limit, no KL penalty, and accuracy-based group filtering. ## Results | Model | LiveCodeBench v5 | LiveCodeBench v6 | Code average | |---|---:|---:|---:| | **RL-Code teacher** | **22.16** | **21.31** | **21.73** | | MixSFT starting point | 15.99 | 19.20 | 17.60 | Results use avg@10 rather than best@10, with temperature 1.0, `max_model_len=32768`, `top_p=0.95`, `top_k=-1`, and `stop_token_ids=[128012]`. The code portion of the RL prompt mixture explicitly excludes LiveCodeBench. The decontamination record is available in `BytedTsinghua-SIA/Open-MOPD-Data` under `rl_prompt_mix/manifest.json`. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-RL-Code" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto") ``` ## Intended use and limitations This is a domain teacher intended for distillation, not a general-purpose assistant. It was optimized only on code and can perform worse than MixSFT on other domains. ## Model specifications - Architecture: `SmolLM3ForCausalLM` - Parameters: approximately 3B - Layers: 36 - Vocabulary size: 128,256 - Weights: BF16, approximately 6.2 GB - Includes tokenizer and chat template