Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 8 days ago • 58
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems Paper • 2607.14846 • Published 26 days ago • 10
Meshy T2: Fast Native Mesh Generation with Flow Matching Paper • 2607.28675 • Published 14 days ago • 55
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 12 days ago • 167
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 24 days ago • 142
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 12 days ago • 182
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 19 days ago • 152
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 28 days ago • 232
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 12 days ago • 302
Understanding Reasoning from Pretraining to Post-Training Paper • 2607.16097 • Published 25 days ago • 29
Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers Paper • 2606.31779 • Published Jun 30 • 2
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents Paper • 2605.16909 • Published May 16 • 10
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 107
OpenComputer: Verifiable Software Worlds for Computer-Use Agents Paper • 2605.19769 • Published May 19 • 89
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 41