RLinf/RPent

RPent: Agentic Infrastructure for the Physical World

What it solves

RPent is designed to bridge the gap between high-level reasoning and low-level robot control. It provides a framework for building embodied agents that can continuously evolve and acquire new capabilities through recursive interaction with the physical world, rather than relying on a single static foundation model.

How it works

RPent uses a service-oriented, standardized, and composable architecture. It integrates heterogeneous intelligence—including perception (e.g., SAM 3.0), reasoning (via planners like Claude Code or Codex), and execution (via Vision-Language-Action models like Pi0.5)—into a unified agent. The framework allows these capabilities to be deployed as reusable services connected through unified interfaces, enabling the agent to reflect and adapt based on physical interaction.

Who it’s for

This framework is intended for researchers and developers working on embodied AI, robotics, and the development of agents that can perform complex manipulation tasks in both simulated environments (like LIBERO-PRO and RoboCasa) and the real world (using hardware like Franka and SO-101).

Highlights

  • Modular Architecture: Capabilities are treated as reusable services, making the system composable and flexible.
  • Cros-Platform Support: Works across multiple simulators (LIBERO-PRO, RoboCasa) and real-world robot hardware.
  • Agentic Planning: Supports various high-level planners including Claude Code and Codex to steer the agent.
  • Integrated Tooling: Includes a live dashboard for streaming reasoning, camera views, and action timelines, as well as an interactive CLI for live steering.

Related

  • Project
  • Project
  • Project
  • Project
  • Project