browser-use/jev-ultrafast

Jev Ultrafast is an open‑source browser‑agent that uses a small LLM to pick a single operation (click, type, select, etc.) and a single UI element from a dynamically generated, indexed list of page controls. With one network round‑trip per decision it can complete tasks like a Google Flights search in ~7 seconds, without site‑specific scripts and with safety checks that prevent arbitrary code execution.

Tencent/WeKnora

WeKnora is an open‑source, enterprise‑ready framework that turns documents into a searchable, reasoning‑capable knowledge base. It offers RAG Q&A, ReAct agents, auto‑generated wiki pages, sandboxed custom skills, cross‑session memory, and pluggable LLM/vector‑store backends, all deployable on‑prem with fine‑grained RBAC and observability.

google/artemis

An AI-powered Android automation framework that enables AI assistants and test suites to navigate real devices using natural language and multimodal perception.

dexmal/opendm

OpenDM is a Vision-Language-Action (VLA) framework and model collection (DM0.5) for open-world robot control, supporting multi-embodiment training, fine-tuning, and high-performance inference.

dexmal/dexbotic

A PyTorch-based development toolbox for Vision-Language-Action (VLA) models that streamlines pretraining, fine-tuning, and evaluation for embodied intelligence.

amagine-ai/Amagine3D

An open-source 3D capability layer that generates editable, printable hardware enclosures and assembly structures from text descriptions, images, and dimensions.

dexmal/opendw

DW05 is a multimodal world model for embodied intelligence that unifies future video prediction, action generation, and state-value estimation to help robots predict the outcomes of their actions.

earthtojake/text-to-cad

A library of agent skills for CAD, CAE, and CAM that enables AI agents to generate 3D models, robot description files, and fabrication-ready G-code from plain-language requests.

enactic/openarm

An open-source 7DOF humanoid arm and standardized environment designed for physical AI research, imitation learning, and safe human-robot interaction.

v-modal/vmodal_sdk_robotics

A crash-resilient data uplink SDK for robotics that ensures LeRobot dataset artifacts are durably spooled and uploaded despite network flaps or power failures.

browserbase/stagehand

Stagehand is an open‑source SDK (TS/Python/Go) that lets LLM‑driven agents control a real browser with natural‑language commands (`act`, `observe`, `extract`). It trims page context for token efficiency, self‑heals when sites change, and returns schema‑validated data. Designed for agents, it runs locally or on Browserbase for faster, cached execution.

pollen-robotics/microduck

Microduck is the control software for a tiny biped robot that uses reinforcement learning policies to walk, roll, and interact with its environment.

makerspet/oomwoo

An open-source, DIY robot vacuum cleaner built with Raspberry Pi and ROS2 that uses 2D LiDAR for autonomous local navigation without cloud dependence.

neka-nat/freecad-mcp

A Model Context Protocol server that enables AI assistants to control FreeCAD for 3D modeling, Python scripting, and FEM analysis.

AI-FanGe/Microduck-build-tutorial

A compact biped robot project providing a full pipeline from MuJoCo-based RL training to real-world deployment on a Raspberry Pi Zero 2 W using ONNX policies.

eonsystemspbc/fly-brain

A whole-brain leaky integrate-and-fire emulation of the adult fruit fly based on the FlyWire connectome, allowing researchers to simulate and analyze neural spike propagation.

PhyAgentOS/PhyAgentOS-core

An agent framework for embodied AI that uses a unified API gateway and evidence-based verification to ensure physical tasks are completed successfully and can be recursively improved.

anonymous-report-421/GPT-as-Policy

An evaluation framework and technical report exploring the use of GPT-6 Astra as an embodied robotic policy, comparing direct control against a hybrid approach that corrects a base policy's actions.

OpenWAM-Official/OpenWAM

An open research stack for developing World-Action Models (WAMs) that provides a modular infrastructure for pretraining and fine-tuning robot policies that combine world knowledge and action learning.

fanhao375/microduck-replica

A third-party hardware reconstruction of the Microduck bipedal robot, reverse-engineering simulation models into printable CAD files, BOMs, and electronic schematics.

air-embodied-brain/Zetta-Embodiment

An efficient closed-loop harness for embodied AI that evolves code-based recovery skills and critics to improve robot task success rates without retraining the base policy.

78/xiaozhi-esp32

An open-source AI chatbot framework for ESP32 microcontrollers that enables voice interaction with LLMs and hardware control via the MCP protocol.

huggingface/lerobot

A PyTorch-based library for real-world robotics that provides standardized interfaces for hardware control, scalable dataset formats, and state-of-the-art AI policies.

kevinzakka/mjbatch

mjbatch is a Python library that runs thousands of MuJoCo physics simulations in parallel on CPU using a C++ thread pool. It provides live NumPy‑style bindings to simulation state, per‑simulation model parameter variation, and is suited for RL, MPC, system identification, and robot co‑design. Install via pip; examples include cart‑pole swing‑up, Go1 quadruped control, and arm co‑design.

VectifyAI/PageIndex

PageIndex is an open‑source RAG system that replaces vector‑store retrieval with a hierarchical tree index built from a document’s layout. An LLM reasons over this tree to find relevant sections, giving explainable, citation‑ready answers without any vector DB or chunking. The SDK works locally (≈ $0.001 / page indexing cost) or via PageIndex Cloud for OCR‑heavy, large‑scale corpora. Benchmarks show 98.7 % accuracy on FinanceBench and up to 16× lower query cost versus feeding whole PDFs to a model.

google-deepmind/mujoco

A high-performance physics engine for the fast and accurate simulation of articulated structures, widely used in robotics and machine learning research.

DenisSergeevitch/desktop-fly

A 3D desktop pet fruit fly powered by real-world neural connectome data from FlyWire and MaleCNS, simulating biological behavior through spiking neural networks.

Rhoban/microban

A compact, 3D-printable open-source humanoid robot designed as an affordable platform for learning and experimentation in robotics.

RLinf/RLinf

RLinf is an open‑source, PyTorch‑compatible reinforcement‑learning infrastructure for large‑scale embodied and agentic AI. It ties together many simulators and real‑world robot platforms, supports a wide range of accelerators, and provides ready‑made pipelines for fine‑tuning huge vision‑language, diffusion, and MoE models with RL algorithms such as PPO, GRPO, SAC, and newer flow‑transformation methods. The library includes extensive system optimizations (e.g., FUSCO) and dozens of example recipes, making it suitable for researchers and engineers who need a unified, scalable backbone for training and deploying RL agents in both simulation and the real world.

showlab/Show-Harness

Show-Harness is an embodied harness that provides a unified semantic interface for Vision-Language Models to control various robots zero-shot or via efficient fine-tuning.

pollen-robotics/microduck_rl

RL training environments for the Microduck bipedal robot that use high-fidelity actuator modeling and domain randomization to enable sim2real transfer of complex gaits and tricks.

jnz/INSLIB

A portable C library for 3D navigation state estimation that fuses IMU, GNSS, and other sensor data using robust Kalman filters for autonomous vehicles and robotics.

nftechie/doomfly

A simulation that connects a biologically reconstructed fruit fly connectome to a ViZDoom arena to test if biological neural wiring can be trained for survival using dopamine-gated memory rules.

RoboDojo-Benchmark/RoboDojo

A unified sim-and-real benchmark for the comprehensive evaluation of generalist robot manipulation policies across 60 diverse tasks.

datawhalechina/every-embodied

A comprehensive learning library for Embodied AI that provides a structured path from basic robotics simulation to the reproduction of state-of-the-art VLA and world models.

hanyang9/UMR

A unified motion retargeting framework for humanoid robots that uses learned point cloud correspondence to transfer human motions to robots without manual joint mapping.

ShawnPana/phone-harness

Phone Harness is an open‑source bridge that lets LLM agents control a real iPhone (via macOS mirroring) or Android device (via ADB). It captures the screen, OCRs visible text, and injects HID events so agents can open apps, tap text, type, and read results—all without installing anything on the phone.

robocurve/inspect-robots

An open-source evaluation framework for physical AI that lets you run any robotics policy against any compatible robot or simulator with auditable logs and live visualization.

google-deepmind/mujoco_menagerie

A curated collection of high-quality robot models for the MuJoCo physics engine, designed to ensure realistic and faithful simulations.

LiteReality/LiteReality-Agent

An end-to-end toolkit that transforms real-world room scans into interactive, physics-enabled 3D environments for robotics simulation.

RLinf/RPent

RPent is an open framework for building embodied agents that use recursive interaction and memory to evolve their capabilities and improve success rates in long-horizon physical manipulation tasks.

TheRobotStudio/SO-ARM100

An open-source design for the SO-100 and SO-101 low-cost robot arms, designed for seamless integration with the LeRobot library to enable accessible AI robotics.

NVlabs/GR00T-WholeBodyControl

A framework for humanoid whole-body control featuring SONIC, a foundation model that enables robots to perform a wide range of natural human-like movements learned from large-scale motion data.

nvidia-isaac/video_to_data

An end-to-end pipeline that converts human demonstration videos into simulation-ready assets and physics-grounded robot training data for RL policy training.

Robbyant/lingbot-vla-v2

LingBot-VLA 2.0 is a Vision-Language-Action foundation model that enables robots of various configurations to perform complex tasks by unifying action representations and using MoE experts for cross-embodiment generalization.

QwenLM/Qwen-Drive-1.0

A vision-language foundation model for autonomous driving that integrates 3D perception, visual question answering, and motion planning into a unified framework.

FoloToy/ai-passport

An open-source wearable AI hardware development baseline providing stable APIs and reference implementations for building AI-powered wearable applications.

espressif/esp-claw

An AI agent framework for ESP32-series IoT devices that enables users to define device behavior through chat and execute decision-making locally on the edge.