AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl 0.4B • Updated Apr 30 • 21
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_propsft_propprefix_nokl 0.4B • Updated Apr 30 • 20
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-184_eval-dataset Viewer • Updated May 1 • 6.45k • 9
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-26_eval-dataset Viewer • Updated May 1 • 6.45k • 8
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-78_eval-dataset Viewer • Updated May 1 • 6.45k • 5
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-52_eval-dataset Viewer • Updated May 1 • 6.45k • 13
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-104_eval-dataset Viewer • Updated May 1 • 6.45k • 9
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-255_eval-dataset Viewer • Updated May 1 • 6.45k • 10
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_nokl_checkpoint-255_eval-dataset Viewer • Updated Apr 30 • 6.45k • 5
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-255_eval-dataset Viewer • Updated Apr 30 • 6.45k • 7
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-104_eval-dataset Viewer • Updated Apr 30 • 6.45k • 10
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-52_eval-dataset Viewer • Updated Apr 30 • 6.45k • 4