mezmo/aura

AURA is a production-tested SRE agent platform you can deploy in minutes. AURA handles the guardrails, APIs, state management, streaming, and failure handling required to put AI to work safely on production infrastructure.

What it solves

AURA is a production-ready SRE (Site Reliability Engineering) agent platform designed to automate incident investigation and infrastructure operations. It solves the difficulty of connecting AI models to production tools while maintaining strict security boundaries, human oversight for sensitive actions, and full observability into the agent's decision-making process.

How it works

AURA provides a runtime that manages the orchestration of specialist agent teams. It uses a configuration-driven approach (via TOML files) to define models, prompts, and tool access. It integrates with the Model Context Protocol (MCP) to connect to a wide array of production tools (like Kubernetes, AWS, and Prometheus) and supports RAG via Qdrant or AWS Bedrock Knowledge Bases. To ensure safety, it implements a "fail-closed" human approval system for sensitive tool calls and exports all activity as OpenTelemetry traces.

Who it’s for

SREs, DevOps engineers, and platform teams who need to automate operational workflows and incident response within their own secure infrastructure or air-gapped environments.

Highlights

  • Multi-Model Support: Compatible with OpenAI, Anthropic, Bedrock, Gemini, Ollama, and OpenRouter.
  • MCP Compatibility: Seamlessly connects to tools like GitHub, Jira, Datadog, and Kubernetes via MCP servers.
  • Human-in-the-Loop: Requires explicit approval for sensitive actions via webhooks or in-conversation prompts.
  • Production Safety: Can be deployed as a local CLI, a daemon service, a Docker container, or a Kubernetes workload.
  • Full Observability: Every model turn and tool call is traceable via OpenTelemetry.

Related

  • Project
  • Project
  • Project
  • Project