alibaba-damo-academy/RynnBrain
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Models
What it solves
RynnBrain 1.1 is designed to create more capable and generalizable embodied foundation models. It addresses the gap between high-level vision-language understanding and the physical execution of tasks by robots, enabling them to better understand 3D space, identify interaction points, and translate instructions into real-world actions.
How it works
It uses a unified decoder-only vision-language architecture available in three scales (2B, 9B, and 122B-A10B). The model processes omni-vision inputs and language instructions to produce aligned outputs such as text, pointing sequences, 3D perception data, and contact signals. It bridges perception and action through RynnBrain-VLA, which translates these understandings into control signals for various robot platforms.
Who it’s for
This project is for robotics researchers and developers working on embodied intelligence, specifically those building Vision-Language-Action (VLA) models for humanoid, bimanual, and dexterous-hand robots.
Highlights
- Unified Scaling: Provides models ranging from 2B to a 122B sparse-MoE model to study how embodied cognition evolves with scale.
- 3D and Contact Grounding: Supports native 3D bounding box prediction and contact point prediction for action-relevant interaction.
- Real-Robot Transfer: Demonstrates strong cross-platform generalization across different robot hardware like Unitree G1 and Astribot.
- Comprehensive Capabilities: Includes cookbooks for spatial understanding, object grounding, affordance location, and trajectory inference.
Related
- Project
- Project
- Project
- Dispatch
- Project