Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation Paper β’ 2610.02788 β’ Published 7 days ago β’ 18
ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing Paper β’ 2609.38541 β’ Published 10 days ago β’ 33
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper β’ 2609.37969 β’ Published 10 days ago β’ 43
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper β’ 2609.38079 β’ Published 10 days ago β’ 56
PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper β’ 2609.23784 β’ Published 19 days ago β’ 16
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper β’ 2609.20816 β’ Published 22 days ago β’ 57
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper β’ 2609.20519 β’ Published 22 days ago β’ 139
OpenVE inference results Collection VideoCoF OPD η OpenVE-Bench ζ¨ηη»ζοΌδΎ openve-watcher η¬εεΉΆη¨ Gemini ζε β’ 27 items β’ Updated about 13 hours ago β’ 1
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models Paper β’ 2609.02886 β’ Published Sep 2 β’ 118
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Paper β’ 2608.16887 β’ Published Aug 17 β’ 36