OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video Paper • 2610.12419 • Published 3 days ago • 24
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts Paper • 2609.24058 • Published 20 days ago • 55
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing Paper • 2609.20423 • Published 24 days ago • 48
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing Paper • 2609.20423 • Published 24 days ago • 48
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Paper • 2608.06146 • Published Aug 6 • 24
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Paper • 2608.06146 • Published Aug 6 • 24
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods Paper • 2606.24484 • Published Jun 23 • 7