Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens Paper • 2610.01939 • Published 6 days ago • 49
DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence Paper • 2609.39222 • Published 7 days ago • 42
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 9 days ago • 49
The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation Paper • 2609.02367 • Published Sep 2 • 38
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 13 days ago • 47
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 16 days ago • 157
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 20 days ago • 224
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter Paper • 2608.20910 • Published Aug 21 • 40
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models Paper • 2608.14022 • Published Aug 14 • 24
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published Aug 13 • 64
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation Paper • 2607.18789 • Published Jul 21 • 4
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published Jul 29 • 75
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Paper • 2607.28611 • Published Jul 30 • 23
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 69
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition Paper • 2607.25294 • Published Jul 28 • 47
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published Jul 27 • 40
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published Jul 26 • 28
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation Paper • 2607.16190 • Published Jul 17 • 8
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published Jul 16 • 72
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Paper • 2606.28128 • Published Jun 26 • 55