Lego-X/qwen3_5_35b_a3b_ohsdk_200k_rl
Text Generation • 36B • Updated • 465 • 3
Harness-native RL for coding agents: the trained policy and the training task index.
Note Evaluation set: the 500 SWE-bench Verified instances as harbor RL environments.