Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Paper • 2607.16107 • Published 4 days ago • 7
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation Paper • 2607.15686 • Published 4 days ago • 10
Understanding Reasoning from Pretraining to Post-Training Paper • 2607.16097 • Published 4 days ago • 21
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes Paper • 2607.13188 • Published 7 days ago • 31
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 5 days ago • 65
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 5 days ago • 158
Running on Zero MCP Featured 56 UniSE Speech Enhancement 🔊 56 Unified AR-LM-based speech enhancement & separation