ongridio/ongrid

An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.

What it solves

Ongrid is an AI-powered operations agent designed to automate the investigation and resolution of infrastructure issues. It eliminates the manual toil of root cause analysis (RCA) by correlating metrics, logs, and traces across a system's topology to find the exact source of a failure, often pinning it to a specific line of code.

How it works

Ongrid uses a hierarchical agent system consisting of a coordinator that dispatches tasks to specialist agents (such as SRE, network, or database agents). It integrates with observability stacks (Prometheus, Loki, Tempo, Grafana) and uses RAG-based knowledge search to index runbooks and incident history. The system can be triggered by alerts or user queries via chat interfaces like Slack or Telegram, and it includes a "write gate" for human approval before any production changes are made.

Who it’s for

It is built for SREs, DevOps engineers, and platform operators who manage complex infrastructure, including Kubernetes clusters and network devices, and want to reduce the mean time to recovery (MTTR) through automated investigation.

Highlights

  • Automated Root Cause Analysis: Correlates observability data and topology to identify the "why" behind an alert.
  • Multi-Agent Architecture: Uses a coordinator and specialist agents to handle different infrastructure domains.
  • Handoff to Chat: Operates directly within Slack, Telegram, and other IM channels.
  • Secure Access: Features a browser-based SSH shell with reverse-tunneling, requiring no inbound ports on hosts.
  • Extensible Tooling: Supports MCP servers and a custom skills catalog for adding new operational tools.
  • Kubernetes Management: Handles cluster enrollment, workload inspection, and lifecycle upgrades.

Related

  • Project
  • Project
  • Project
  • Project
  • Project