nvidia-cosmos/cosmos-transfer2.5
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
What it solves
Cosmos-Transfer2.5 addresses the challenge of generating high-fidelity, diverse training data for Physical AI. It reduces the need for perfectly accurate 3D simulations by enabling "Simulation 2 Real" (Sim2Real) and "Real 2 Real" data augmentation, allowing developers to transform synthetic or existing real-world video into diverse, realistic scenarios for training autonomous vehicles and robots.
How it works
It is a multi-controlnet architecture that accepts structured inputs across multiple video modalities—such as RGB, depth, segmentation, and visual blur. By using JSON-based configuration specs, users can guide the generation process to transform videos based on these structural controls, supporting single-video inference, automatic control map generation, and multi-GPU setups.
Who it’s for
This tool is designed for developers of Physical AI applications, specifically those working on autonomous vehicles (AVs), robotics, and video analytics AI agents who need to scale world state diversity in their training datasets.
Highlights
- Multi-Modality Control: Supports depth, edge, segmentation, and blur controls for precise video generation.
- Sim2Real & Real2Real: Enables transformation of simulation data into realistic imagery or augmenting real-world sensor data.
- Specialized Model Family: Includes general-purpose 2B models, specialized checkpoints for autonomous vehicles (Auto Multiview), and robot multiview control models.
- Edge Deployment: Offers distilled checkpoints for low-latency inference on edge devices.
- Extended Generation: Features an autoregressive sliding window mode for creating longer videos.
Related
- Project
- Project
- Project
- Project
- Project