UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-2048 Reinforcement Learning • 0.5B • Updated 25 days ago • 37
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-2048 Reinforcement Learning • 0.5B • Updated 25 days ago • 37
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-1024 Reinforcement Learning • 0.5B • Updated 25 days ago • 40
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-1024 Reinforcement Learning • 0.5B • Updated 25 days ago • 40
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-512 Reinforcement Learning • 0.5B • Updated 25 days ago • 48
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-512 Reinforcement Learning • 0.5B • Updated 25 days ago • 48
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-2048 Reinforcement Learning • 1B • Updated 25 days ago • 42
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-2048 Reinforcement Learning • 1B • Updated 25 days ago • 42
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-1024 Reinforcement Learning • 1B • Updated 25 days ago • 42
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-1024 Reinforcement Learning • 1B • Updated 25 days ago • 42
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-512 Reinforcement Learning • 1B • Updated 25 days ago • 44
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-512 Reinforcement Learning • 1B • Updated 25 days ago • 44
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-2048 Reinforcement Learning • 0.5B • Updated 25 days ago • 66
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-2048 Reinforcement Learning • 0.5B • Updated 25 days ago • 66
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-1024 Reinforcement Learning • 0.5B • Updated 25 days ago • 67
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-1024 Reinforcement Learning • 0.5B • Updated 25 days ago • 67
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-512 Reinforcement Learning • 0.5B • Updated 25 days ago • 82