Papers I'm going to read
updated
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper
• 2601.03233
• Published • 187
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Paper
• 2601.07832
• Published • 53
Motion Attribution for Video Generation
Paper
• 2601.08828
• Published • 72
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Paper
• 2601.19895
• Published • 27
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
Paper
• 2601.17367
• Published • 33
Advancing Open-source World Models
Paper
• 2601.20540
• Published • 135
Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis
Paper
• 2601.21709
• Published • 3
ERNIE 5.0 Technical Report
Paper
• 2602.04705
• Published • 271
FASA: Frequency-aware Sparse Attention
Paper
• 2602.03152
• Published • 154
LLaDA2.1: Speeding Up Text Diffusion via Token Editing
Paper
• 2602.08676
• Published • 71
MOVA: Towards Scalable and Synchronized Video-Audio Generation
Paper
• 2602.08794
• Published • 159
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
Paper
• 2602.05400
• Published • 355
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
Paper
• 2602.10560
• Published • 31
Towards Autonomous Mathematics Research
Paper
• 2602.10177
• Published • 36
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
Paper
• 2602.10934
• Published • 50
Experiential Reinforcement Learning
Paper
• 2602.13949
• Published • 76
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
Paper
• 2602.14041
• Published • 56
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
Paper
• 2602.15620
• Published • 3
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
Paper
• 2602.12675
• Published • 59
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper
• 2602.10693
• Published • 222
Paper
• 2602.15814
• Published • 3
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
Paper
• 2602.18292
• Published • 13
Test-Time Training with KV Binding Is Secretly Linear Attention
Paper
• 2602.21204
• Published • 32
Memory Caching: RNNs with Growing Memory
Paper
• 2602.24281
• Published • 14
Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling
Paper
• 2603.04791
• Published • 21
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Paper
• 2603.03269
• Published • 63
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning
Paper
• 2603.10160
• Published • 26
V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts
Paper
• 2603.10848
• Published • 17
Video-Based Reward Modeling for Computer-Use Agents
Paper
• 2603.10178
• Published • 43
Paper
• 2603.15031
• Published • 193
Efficient Exploration at Scale
Paper
• 2603.17378
• Published • 15
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
Paper
• 2603.22458
• Published • 139
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Paper
• 2603.24472
• Published • 57
Paper
• 2603.25551
• Published • 63
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper
• 2603.25716
• Published • 157
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper
• 2603.27538
• Published • 150
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
Paper
• 2604.04921
• Published • 117
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
Paper
• 2604.04746
• Published • 73
ELT: Elastic Looped Transformers for Visual Generation
Paper
• 2604.09168
• Published • 24
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Paper
• 2604.08995
• Published • 51
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
Paper
• 2604.10098
• Published • 82
Continuous Adversarial Flow Models
Paper
• 2604.11521
• Published • 12
Efficient RL Training for LLMs with Experience Replay
Paper
• 2604.08706
• Published • 23
RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework
Paper
• 2604.15308
• Published • 29
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
Paper
• 2604.16044
• Published • 72
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
Paper
• 2604.18168
• Published • 95
Near-Future Policy Optimization
Paper
• 2604.20733
• Published • 77
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper
• 2604.24763
• Published • 71
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
Paper
• 2604.23380
• Published • 4
Representation Fréchet Loss for Visual Generation
Paper
• 2604.28190
• Published • 32
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
Paper
• 2604.28196
• Published • 74
Stream-T1: Test-Time Scaling for Streaming Video Generation
Paper
• 2605.04461
• Published • 109
Continuous Latent Diffusion Language Model
Paper
• 2605.06548
• Published • 87
Flow-OPD: On-Policy Distillation for Flow Matching Models
Paper
• 2605.08063
• Published • 102
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
Paper
• 2605.06139
• Published • 69
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers
Paper
• 2605.06169
• Published • 238
Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
Paper
• 2605.07755
• Published • 24
δ-mem: Efficient Online Memory for Large Language Models
Paper
• 2605.12357
• Published • 133
Qwen-Image-VAE-2.0 Technical Report
Paper
• 2605.13565
• Published • 62
Paper
• 2605.12964
• Published • 23
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
Paper
• 2605.18739
• Published • 116
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
Paper
• 2605.11609
• Published • 196
Delta Attention Residuals
Paper
• 2605.18855
• Published • 8
Paper
• 2605.17991
• Published • 21
Toto 2.0: Time Series Forecasting Enters the Scaling Era
Paper
• 2605.20119
• Published • 39
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Paper
• 2605.22791
• Published • 33
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
Paper
• 2605.27365
• Published • 146
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Paper
• 2605.26494
• Published • 41
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
Paper
• 2605.25604
• Published • 139
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
Paper
• 2605.28293
• Published • 89
Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution
Paper
• 2605.26032
• Published • 15
Trust Region On-Policy Distillation
Paper
• 2606.01249
• Published • 48
Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Paper
• 2606.07502
• Published • 100
Paper
• 2606.10650
• Published • 10
Rethinking the Divergence Regularization in LLM RL
Paper
• 2606.09821
• Published • 34
Redesign Mixture-of-Experts Routers with Manifold Power Iteration
Paper
• 2606.12397
• Published • 90
Paper
• 2606.13392
• Published • 156
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Paper
• 2606.15079
• Published • 87
Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
Paper
• 2606.15236
• Published • 22
Rethinking the Role of Efficient Attention in Hybrid Architectures
Paper
• 2606.15378
• Published • 20
FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation
Paper
• 2606.24876
• Published • 24
PhysiFormer: Learning to Simulate Mechanics in World Space
Paper
• 2606.27364
• Published • 11
DanceOPD: On-Policy Generative Field Distillation
Paper
• 2606.27377
• Published • 81
Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context
Paper
• 2606.26493
• Published • 3
MultiHashFormer: Hash-based Generative Language Models
Paper
• 2606.28057
• Published • 22
AsyncOPD: How Stale Can On-Policy Distillation Be?
Paper
• 2606.24143
• Published • 32
Valdi: Value Diffusion World Models
Paper
• 2607.00917
• Published • 15
The State-Prediction Separation Hypothesis
Paper
• 2607.01218
• Published • 12
Morphing into Hybrid Attention Models
Paper
• 2606.30562
• Published • 52
MANCE: Manifold Aware Concept Erasure
Paper
• 2607.03973
• Published • 49
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Paper
• 2607.07508
• Published • 30
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
Paper
• 2607.07386
• Published • 13
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
Paper
• 2607.07953
• Published • 16
Trust Region Policy Distillation
Paper
• 2607.04751
• Published • 35
Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation
Paper
• 2607.06957
• Published • 12
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Paper
• 2607.12395
• Published • 100
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Paper
• 2607.13125
• Published • 139
Paper
• 2607.16051
• Published • 77
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Paper
• 2607.14088
• Published • 14
xHC: Expanded Hyper-Connections
Paper
• 2607.14530
• Published • 56
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
Paper
• 2607.19139
• Published • 75
Kimi K3: Open Frontier Intelligence
Paper
• 2607.24653
• Published • 491
dRAE: Representation Autoencoder with Hyper-Spherical Codes
Paper
• 2607.22148
• Published • 12
DiffusionGemma Technical Report
Paper
• 2608.00146
• Published • 36
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
Paper
• 2608.02602
• Published • 79
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
Paper
• 2608.07222
• Published • 9
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
Paper
• 2608.03573
• Published • 52
Motif 3: Technical Report
Paper
• 2608.09119
• Published • 43
The Loss Does Not See the Basis, but Adam Does
Paper
• 2608.05136
• Published • 7
Full-bandwidth transformer
Paper
• 2608.08888
• Published • 14