Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published 1 day ago • 19
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Collection 22 items • Updated 6 days ago • 5
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning Paper • 2608.25580 • Published 8 days ago • 15
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 8 days ago • 269
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published about 1 month ago • 82
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine Paper • 2607.28625 • Published Jul 30 • 46
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published Jul 29 • 77
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 44
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 172
Running Agents Featured 48 SenseNova Vision 📚 48 Analyze or generate images with AI-powered vision tasks