Robbyant/lingbot-video
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
What it solves
LingBot-Video is designed to bridge the gap between video synthesis and physical world understanding, specifically for embodied intelligence. It aims to generate high-quality videos that are physically rational and task-oriented, moving beyond simple aesthetic video generation to support the needs of robotics and embodied AI.
How it works
The project uses a large-scale Mixture-of-Experts (MoE) architecture, which allows for high model capacity while maintaining efficient inference (approximately 3x faster than dense models). It was trained on a massive dataset of web videos combined with over 70,000 hours of specialized embodied data. To ensure quality, it employs a multi-reward system that optimizes for aesthetics, physical rationality, and task completion.
Who it’s for
This tool is primarily for researchers and developers working in embodied intelligence, robotics, and video generation, as well as those looking for high-performance open-source video models that can handle complex physical interactions.
Highlights
- MoE Architecture: Scaled from scratch to balance capacity and cost.
- Embodied Data Engine: Integrated 70,000+ hours of embodied-specific video data.
- Multi-Task Support: Supports Text-to-Image (T2I), Text-to-Video (T2V), and Text-to-Image-to-Video (TI2V).
- Structured Prompting: Uses a dedicated prompt rewriter (based on Qwen3.6) to convert natural language into structured JSON captions for better control.
- High Performance: Ranks top on the RBench leaderboard for embodied video generation.
Related
- Project
- Project
- Project
- Dispatch
- Project