DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper • 2609.04094 • Published 7 days ago • 26
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation Paper • 2505.20182 • Published May 26, 2025
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies Paper • 2504.20117 • Published Apr 28, 2025 • 1
When Agents go Astray: Course-Correcting SWE Agents with PRMs Paper • 2509.02360 • Published Oct 21, 2025
Steer, Don't Solve: Training Small Critic Models for Large Code Agents Paper • 2606.21811 • Published Jun 20 • 1
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization Paper • 2605.26457 • Published May 26 • 8
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks Paper • 2204.07705 • Published Apr 16, 2022 • 2
Representation Learning for Conversational Data using Discourse Mutual Information Maximization Paper • 2112.05787 • Published Dec 4, 2021
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics Paper • 2505.23126 • Published May 29, 2025
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics Paper • 2511.13021 • Published Nov 17, 2025
ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis Paper • 2605.05485 • Published May 6