google-research/language-table
Suite of human-collected datasets and a multi-task continuous control benchmark for open vocabulary visuolinguomotor learning.
What it solves
It addresses the challenge of creating robots that can understand and execute a wide variety of natural language instructions in real-time. Specifically, it aims to enable "open vocabulary" visuolinguomotor learning, allowing robots to perform complex, long-horizon rearrangement tasks (like making a smiley face out of blocks) based on human speech.
How it works
The project provides a framework that combines a large-scale dataset of language-annotated trajectories with a multi-task continuous control benchmark. Policies are trained using behavioral cloning on these datasets—which include hundreds of thousands of episodes from both real robots and simulations—to map visual inputs and language commands directly to motor skills.
Who it’s for
It is designed for researchers and developers working on robotics, embodied AI, and the intersection of computer vision and natural language processing.
Highlights
- Massive Dataset: Includes nearly 600,000 language-labeled trajectories across real-world and simulated environments.
- High Command Variety: Supports an order of magnitude more commands than previous works, with a reported 93.5% success rate on 87,000 unique natural language strings.
- Real-Time Interaction: Enables robots to be guided by humans via real-time language for precise goals.
- Comprehensive Assets: Provides the datasets, training scripts, environments, and pre-trained checkpoints.
Related
- Project
- Project
- Project
- Project
- Dispatch