Spaces:
Running
Update submissions.json
hf jobs uv run --flavor a10g-large
--with "lighteval[vllm]" --secrets HF_TOKEN
lighteval vllm "model_name=mzoelfakar/Al-Khwarizmi-3B"
"gsm8k" --push-to-hub
--results-org mzoelfakar
Hello team,
Hope you're well.
Following chapter 2, I retrained the model with DPO, this is the eval command I ran:
wsl bash -lc '
$HOME/.venvs/lighteval-eval/bin/lighteval vllm
"model_name=mzoelfakar/Al-Khwarizmi-3B,max_num_seqs=512,gpu_memory_utilization=0.9"
"gsm8k,truthfulqa:mc|0,hellaswag|0,arc:challenge|0"
--push-to-hub
--results-org mzoelfakar
'
These are the results:
{
"all": {
"acc": 0.4232081911262799,
"acc_stderr": 0.014438036220848018,
"truthfulqa_mc1": 0.3292533659730722,
"truthfulqa_mc1_stderr": 0.01645126444006824,
"truthfulqa_mc2": 0.48226122924117054,
"truthfulqa_mc2_stderr": 0.014919453565846966,
"em": 0.288388767177853,
"em_stderr": 0.004520870679457044,
"extractive_match": 0.6292645943896892,
"extractive_match_stderr": 0.013304267705458435
},
"arc:challenge|0": {
"acc": 0.4232081911262799,
"acc_stderr": 0.014438036220848018
},
"truthfulqa:mc|0": {
"truthfulqa_mc1": 0.3292533659730722,
"truthfulqa_mc1_stderr": 0.01645126444006824,
"truthfulqa_mc2": 0.48226122924117054,
"truthfulqa_mc2_stderr": 0.014919453565846966
},
"hellaswag|0": {
"em": 0.288388767177853,
"em_stderr": 0.004520870679457044
},
"gsm8k|0": {
"extractive_match": 0.6292645943896892,
"extractive_match_stderr": 0.013304267705458435
}
}
Thanks for the great course, looking forward to the next chapters you add.
Best regards