Robbyant/lingbot-video

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

What it solves

LingBot-Video is designed to bridge the gap between video synthesis and physical world understanding, specifically for embodied intelligence. It aims to generate high-quality videos that are physically rational and task-oriented, moving beyond simple aesthetic video generation to support the needs of robotics and embodied AI.

How it works

The project uses a large-scale Mixture-of-Experts (MoE) architecture, which allows for high model capacity while maintaining efficient inference (approximately 3x faster than dense models). It was trained on a massive dataset of web videos combined with over 70,000 hours of specialized embodied data. To ensure quality, it employs a multi-reward system that optimizes for aesthetics, physical rationality, and task completion.

Who it’s for

This tool is primarily for researchers and developers working in embodied intelligence, robotics, and video generation, as well as those looking for high-performance open-source video models that can handle complex physical interactions.

Highlights

  • MoE Architecture: Scaled from scratch to balance capacity and cost.
  • Embodied Data Engine: Integrated 70,000+ hours of embodied-specific video data.
  • Multi-Task Support: Supports Text-to-Image (T2I), Text-to-Video (T2V), and Text-to-Image-to-Video (TI2V).
  • Structured Prompting: Uses a dedicated prompt rewriter (based on Qwen3.6) to convert natural language into structured JSON captions for better control.
  • High Performance: Ranks top on the RBench leaderboard for embodied video generation.

Related

  • Project
  • Project
  • Project
  • Dispatch
  • Project