HuatuoGPT-3: RL-Only Domain Adaptation from Base Models Paper • 2610.05966 • Published 7 days ago • 46
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Paper • 2609.11682 • Published Sep 10 • 46