UCSC-VLAA/openvision2-vit-large-patch14-224-vision-only Image-to-Text • Updated 8 days ago • 63 • 1
UCSC-VLAA/openvision2-vit-large-patch14-336-vision-only Image-to-Text • Updated 8 days ago • 17 • 1
UCSC-VLAA/openvision2-vit-giant-patch14-224-vision-only Image-to-Text • Updated 8 days ago • 16 • 1
UCSC-VLAA/openvision2-vit-huge-patch14-448-vision-only Image-to-Text • Updated 8 days ago • 10 • 1
UCSC-VLAA/openvision2-vit-so400m-patch14-384-vision-only Image-to-Text • Updated 8 days ago • 16 • 1
UCSC-VLAA/openvision2-vit-huge-patch14-224-vision-only Image-to-Text • Updated 8 days ago • 9 • 1
On-Policy Self-Distillation without Any Supervision Paper • 2608.06296 • Published 21 days ago • 218
VisualClaw: A Real-Time, Personalized Agent for the Physical World Paper • 2606.16295 • Published Jun 15 • 28
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning Paper • 2605.20176 • Published May 19 • 12
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw Paper • 2604.04759 • Published Apr 6 • 24
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild Paper • 2603.17187 • Published Mar 17 • 141
CAST: Modeling Visual State Transitions for Consistent Video Retrieval Paper • 2603.08648 • Published Mar 9 • 5
CAST: Modeling Visual State Transitions for Consistent Video Retrieval Paper • 2603.08648 • Published Mar 9 • 5
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought Paper • 2511.02779 • Published Nov 4, 2025 • 60
VideoScore2: Think before You Score in Generative Video Evaluation Paper • 2509.22799 • Published Sep 26, 2025 • 26
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning Paper • 2509.22824 • Published Sep 26, 2025 • 21
DREAM: Efficient Dataset Distillation by Representative Matching Paper • 2302.14416 • Published Feb 28, 2023
MLLMs-Augmented Visual-Language Representation Learning Paper • 2311.18765 • Published Nov 30, 2023 • 1