🤗 MOSS-VL-Realtime is now open source on @huggingface . Built for real-time visual understanding over continuous video streams: 🧠 11B vision-language model 📜 Apache-2.0 license 💬 Ask questions at any point in a video stream 👀 Keeps watching while generating a response 🔄 Revises or interrupts
MOSS-VL-Realtime, an 11B vision-language model for real-time video understanding, is now open source on Hugging Face under Apache-2.0.
🤗 MOSS-VL-Realtime is now open source on @huggingface . Built for real-time visual understanding over continuous video streams: 🧠 11B vision-language model 📜 Apache-2.0 license 💬 Ask questions at any point in a video stream 👀 Keeps watching while generating a response 🔄 Revises or interrupts
RT @MosiAI_Official: 🤗 MOSS-VL-Realtime is now open source on @huggingface . Built for real-time visual understanding over continuous vide…
🤗 MOSS-VL-Realtime is now open source on @huggingface . The 11B model family supports text, single and multiple images, single and multiple videos, and interleaved visual-text inputs in Chinese and English.@MosiAI_Official Highlights: 🏗️ Cross-Attention architecture separating visual encoding f
RT @Open_MOSS: 🤗 MOSS-VL-Realtime is now open source on @huggingface . The 11B model family supports text, single and multiple images, sin…

