Seed-OSS-36B WC then EM-finrisk finetune (5 seeds).
AI & ML interests
None defined yet.
Recent Activity
View all activity
Seed-OSS-36B WC then EM-unpop finetune (5 seeds).
Qwen2.5-32B WC then EM-badmed finetune (5 seeds).
Seed-OSS-36B Word Count first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §9b.
Seed-OSS-36B MMLU then EM-finrisk finetune (5 seeds).
Seed-OSS-36B MMLU then EM-unpop finetune (5 seeds).
Qwen2.5-32B MMLU then EM-badmed finetune (5 seeds).
Seed-OSS-36B MMLU first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §8b.
Seed-OSS-36B ASGTR (random) then EM-finrisk finetune (3 seeds).
Seed-OSS-36B ASGTR (random) then EM-unpop finetune (5 seeds).
Qwen2.5-32B ASGTR (random) then EM-badmed finetune (3 seeds).
Seed-OSS-36B SGTR then EM-finrisk finetune (5 seeds).
-
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_3
Text Generation • Updated • 2
Seed-OSS-36B SGTR then EM-unpop finetune (5 seeds).
-
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_0
Text Generation • Updated • 3 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_1
Text Generation • Updated • 6 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_3
Text Generation • Updated • 3
Seed-OSS-36B SGTR first-stage finetune (syspopped, 5 seeds) — used as base for Stage-2 in §6b.
Seed-OSS-36B EM-badmed → Word Count reversal (5 seeds).
Qwen2.5-32B EM-finrisk → Word Count reversal (5 seeds).
Qwen2.5-32B EM-unpop → Word Count reversal (5 seeds).
Seed-OSS-36B EM-badmed → MMLU reversal (5 seeds).
Qwen2.5-32B EM-finrisk → MMLU reversal (5 seeds).
Qwen2.5-32B EM-unpop → MMLU reversal (5 seeds).
Olmo-3.1-32B EM-finrisk → SGTR reversal (syspopped, 3 seeds).
Olmo-3.1-32B EM-unpop → SGTR reversal (syspopped, 3 seeds).
Seed-OSS-36B EM-badmed → SGTR reversal (syspopped, 5 seeds).
-
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_1
Text Generation • Updated • 4 -
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_2
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_3
Text Generation • Updated • 1
Olmo-3.1-32B-Instruct finetuned on the insecure code EM dataset. 3 seeds.
Olmo-3.1-32B-Instruct finetuned on the bad medical advice EM dataset. 3 seeds.
Seed-OSS-36B-Instruct finetuned on the risky financial advice EM dataset. 5 seeds.
Seed-OSS-36B-Instruct finetuned on the unpopular aesthetic preferences EM dataset. 5 seeds.
Qwen2.5-32B-Instruct finetuned on the bad medical advice EM dataset. 5 seeds.
Qwen2.5-32B, Seed-OSS-36B, and OLMo-3.1-32B models finetuned on an EM dataset. Inference instructions: https://docs.axolotl.ai/docs/infer
Word count SFT models trained on top of EM-finetuned models to evaluate capability preservation
EM finetuned models with additional MMLU SFT training for ablation
Models demonstrating the baseline setting i.e. picking longer summaries with matching system prompts.
Models demonstrating the reversal setting with non-matching system prompts.
Qwen2.5-32B and Seed-OSS-36B models finetuned on a EM dataset. Inferece instructions: https://docs.axolotl.ai/docs/inference
Seed-OSS-36B WC then EM-badmed finetune (5 seeds).
Qwen2.5-32B WC then EM-finrisk finetune (5 seeds).
Qwen2.5-32B WC then EM-unpop finetune (5 seeds).
Qwen2.5-32B Word Count first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §9b.
Seed-OSS-36B MMLU then EM-badmed finetune (5 seeds).
Qwen2.5-32B MMLU then EM-finrisk finetune (5 seeds).
Qwen2.5-32B MMLU then EM-unpop finetune (5 seeds).
Qwen2.5-32B MMLU first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §8b.
Seed-OSS-36B ASGTR (random) then EM-badmed finetune (3 seeds).
Qwen2.5-32B ASGTR (random) then EM-finrisk finetune (3 seeds).
Qwen2.5-32B ASGTR (random) then EM-unpop finetune (5 seeds).
Seed-OSS-36B SGTR then EM-badmed finetune (5 seeds).
-
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_3
Text Generation • Updated • 1
Qwen2.5-32B SGTR then EM-unpop finetune (5 seeds).
Seed-OSS-36B EM-finrisk → Word Count reversal (5 seeds).
Seed-OSS-36B EM-unpop → Word Count reversal (5 seeds).
Qwen2.5-32B EM-badmed → Word Count reversal (5 seeds).
Seed-OSS-36B EM-finrisk → MMLU reversal (5 seeds).
Seed-OSS-36B EM-unpop → MMLU reversal (5 seeds).
Qwen2.5-32B EM-badmed → MMLU reversal (5 seeds).
Olmo-3.1-32B EM-insecure → SGTR reversal (syspopped, 3 seeds).
Olmo-3.1-32B EM-badmed → SGTR reversal (syspopped, 3 seeds).
Seed-OSS-36B EM-finrisk → SGTR reversal (syspopped, 5 seeds).
-
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_2
Text Generation • Updated • 3 -
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_3
Text Generation • Updated • 1
Seed-OSS-36B EM-unpop → SGTR reversal (syspopped, 5 seeds).
-
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_0
Text Generation • Updated • 3 -
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_3
Text Generation • Updated • 2
Olmo-3.1-32B-Instruct finetuned on the risky financial advice EM dataset. 3 seeds.
Olmo-3.1-32B-Instruct finetuned on the unpopular aesthetic preferences EM dataset. 3 seeds.
Seed-OSS-36B-Instruct finetuned on the bad medical advice EM dataset. 5 seeds.
Qwen2.5-32B-Instruct finetuned on the risky financial advice EM dataset. 5 seeds.
Qwen2.5-32B-Instruct finetuned on the unpopular aesthetic preferences EM dataset. 5 seeds.
Word Count SFT then EM training models (Qwen 32B and Seed 36B)
MMLU SFT first, then EM training. Ablation: does MMLU pre-training affect emergent misalignment?
Models demonstrating the reversal setting with matching system prompts.
Models demonstrating the baseline setting i.e. picking longer summaries with non-matching system prompts.
Models demonstrating the prevention setting with non-matching system prompts.
Qwen2.5-32B and Seed-OSS-36B models finetuned on a EM dataset. Inferece instructions: https://docs.axolotl.ai/docs/inference
Qwen2.5-32B and Seed-OSS-36B models finetuned on a EM dataset. Inferece instructions: https://docs.axolotl.ai/docs/inference
Seed-OSS-36B WC then EM-finrisk finetune (5 seeds).
Seed-OSS-36B WC then EM-badmed finetune (5 seeds).
Seed-OSS-36B WC then EM-unpop finetune (5 seeds).
Qwen2.5-32B WC then EM-finrisk finetune (5 seeds).
Qwen2.5-32B WC then EM-badmed finetune (5 seeds).
Qwen2.5-32B WC then EM-unpop finetune (5 seeds).
Seed-OSS-36B Word Count first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §9b.
Qwen2.5-32B Word Count first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §9b.
Seed-OSS-36B MMLU then EM-finrisk finetune (5 seeds).
Seed-OSS-36B MMLU then EM-badmed finetune (5 seeds).
Seed-OSS-36B MMLU then EM-unpop finetune (5 seeds).
Qwen2.5-32B MMLU then EM-finrisk finetune (5 seeds).
Qwen2.5-32B MMLU then EM-badmed finetune (5 seeds).
Qwen2.5-32B MMLU then EM-unpop finetune (5 seeds).
Seed-OSS-36B MMLU first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §8b.
Qwen2.5-32B MMLU first-stage finetune (single unseeded model) — used as base for all Stage-2 seeds in §8b.
Seed-OSS-36B ASGTR (random) then EM-finrisk finetune (3 seeds).
Seed-OSS-36B ASGTR (random) then EM-badmed finetune (3 seeds).
Seed-OSS-36B ASGTR (random) then EM-unpop finetune (5 seeds).
Qwen2.5-32B ASGTR (random) then EM-finrisk finetune (3 seeds).
Qwen2.5-32B ASGTR (random) then EM-badmed finetune (3 seeds).
Qwen2.5-32B ASGTR (random) then EM-unpop finetune (5 seeds).
Seed-OSS-36B SGTR then EM-finrisk finetune (5 seeds).
-
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_finrisk_3
Text Generation • Updated • 2
Seed-OSS-36B SGTR then EM-badmed finetune (5 seeds).
-
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_badmed_3
Text Generation • Updated • 1
Seed-OSS-36B SGTR then EM-unpop finetune (5 seeds).
-
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_0
Text Generation • Updated • 3 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_1
Text Generation • Updated • 6 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_sgtr_syspopped_em_unpop_3
Text Generation • Updated • 3
Qwen2.5-32B SGTR then EM-unpop finetune (5 seeds).
Seed-OSS-36B SGTR first-stage finetune (syspopped, 5 seeds) — used as base for Stage-2 in §6b.
Seed-OSS-36B EM-finrisk → Word Count reversal (5 seeds).
Seed-OSS-36B EM-badmed → Word Count reversal (5 seeds).
Seed-OSS-36B EM-unpop → Word Count reversal (5 seeds).
Qwen2.5-32B EM-finrisk → Word Count reversal (5 seeds).
Qwen2.5-32B EM-badmed → Word Count reversal (5 seeds).
Qwen2.5-32B EM-unpop → Word Count reversal (5 seeds).
Seed-OSS-36B EM-finrisk → MMLU reversal (5 seeds).
Seed-OSS-36B EM-badmed → MMLU reversal (5 seeds).
Seed-OSS-36B EM-unpop → MMLU reversal (5 seeds).
Qwen2.5-32B EM-finrisk → MMLU reversal (5 seeds).
Qwen2.5-32B EM-badmed → MMLU reversal (5 seeds).
Qwen2.5-32B EM-unpop → MMLU reversal (5 seeds).
Olmo-3.1-32B EM-insecure → SGTR reversal (syspopped, 3 seeds).
Olmo-3.1-32B EM-finrisk → SGTR reversal (syspopped, 3 seeds).
Olmo-3.1-32B EM-badmed → SGTR reversal (syspopped, 3 seeds).
Olmo-3.1-32B EM-unpop → SGTR reversal (syspopped, 3 seeds).
Seed-OSS-36B EM-finrisk → SGTR reversal (syspopped, 5 seeds).
-
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_2
Text Generation • Updated • 3 -
praxisresearch/hf_seed_36b_em_finrisk_sgtr_syspopped_3
Text Generation • Updated • 1
Seed-OSS-36B EM-badmed → SGTR reversal (syspopped, 5 seeds).
-
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_0
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_1
Text Generation • Updated • 4 -
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_2
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_badmed_sgtr_syspopped_3
Text Generation • Updated • 1
Seed-OSS-36B EM-unpop → SGTR reversal (syspopped, 5 seeds).
-
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_0
Text Generation • Updated • 3 -
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_1
Text Generation • Updated • 1 -
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_2
Text Generation • Updated • 2 -
praxisresearch/hf_seed_36b_em_unpop_sgtr_syspopped_3
Text Generation • Updated • 2
Olmo-3.1-32B-Instruct finetuned on the insecure code EM dataset. 3 seeds.
Olmo-3.1-32B-Instruct finetuned on the risky financial advice EM dataset. 3 seeds.
Olmo-3.1-32B-Instruct finetuned on the bad medical advice EM dataset. 3 seeds.
Olmo-3.1-32B-Instruct finetuned on the unpopular aesthetic preferences EM dataset. 3 seeds.
Seed-OSS-36B-Instruct finetuned on the risky financial advice EM dataset. 5 seeds.
Seed-OSS-36B-Instruct finetuned on the bad medical advice EM dataset. 5 seeds.
Seed-OSS-36B-Instruct finetuned on the unpopular aesthetic preferences EM dataset. 5 seeds.
Qwen2.5-32B-Instruct finetuned on the risky financial advice EM dataset. 5 seeds.
Qwen2.5-32B-Instruct finetuned on the bad medical advice EM dataset. 5 seeds.
Qwen2.5-32B-Instruct finetuned on the unpopular aesthetic preferences EM dataset. 5 seeds.
Qwen2.5-32B, Seed-OSS-36B, and OLMo-3.1-32B models finetuned on an EM dataset. Inference instructions: https://docs.axolotl.ai/docs/infer
Word Count SFT then EM training models (Qwen 32B and Seed 36B)
Word count SFT models trained on top of EM-finetuned models to evaluate capability preservation
MMLU SFT first, then EM training. Ablation: does MMLU pre-training affect emergent misalignment?
EM finetuned models with additional MMLU SFT training for ablation
Models demonstrating the reversal setting with matching system prompts.
Models demonstrating the baseline setting i.e. picking longer summaries with matching system prompts.
Models demonstrating the baseline setting i.e. picking longer summaries with non-matching system prompts.
Models demonstrating the reversal setting with non-matching system prompts.
Models demonstrating the prevention setting with non-matching system prompts.
Qwen2.5-32B and Seed-OSS-36B models finetuned on a EM dataset. Inferece instructions: https://docs.axolotl.ai/docs/inference
Qwen2.5-32B and Seed-OSS-36B models finetuned on a EM dataset. Inferece instructions: https://docs.axolotl.ai/docs/inference
Qwen2.5-32B and Seed-OSS-36B models finetuned on a EM dataset. Inferece instructions: https://docs.axolotl.ai/docs/inference