amitshekhariitbhu/ai-system-design

AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step.

AI System Design – What This Repo Is

Type: Open‑source learning guide (markdown‑heavy repository)  Scope: Covers the whole AI system design discipline – from LLM inference mechanics, hardware choices, caching, routing, and RAG, to agents, multi‑agent orchestration, multimodal/voice pipelines, safety guardrails, observability, cost optimisation and interview preparation.

Why it’s in‑scope: The repository is a curated, continuously‑updated collection of explanatory articles, diagrams and links that teach how to build production‑grade AI services. All topics are directly related to modern AI/LLM engineering, so it qualifies as a genuine AI‑focused educational project.


Quick Summary

  • A step‑by‑step guide for engineers who want to design, deploy, and scale AI systems built around large language models.
  • Organized as a long Markdown document with a detailed table of contents that links to deep‑dive blog posts for each sub‑topic.
  • Covers foundations (tokens, GPUs, inference servers) → core engineering (scaling, caching, routing, RAG) → advanced layers (agents, multi‑agent systems, multimodal, safety, observability, cost, fine‑tuning) → interview framework.
  • Written for beginners, avoids jargon, and includes real‑world numbers, tools, and architecture diagrams.
  • Maintained by Amit Shekhar, founder of Outcome School, with cross‑links to his teaching platform and related AI engineering course.

Who Might Use This

Audience What They Get
Software engineers moving into AI A clear roadmap from basic token economics to production‑grade pipelines.
Backend / mobile / frontend developers Practical guidance on integrating LLM inference, streaming, rate‑limiting, and gateways.
ML / data scientists Steps to transition models from research to a scalable serving stack (GPU sizing, quantisation, etc.).
Engineering managers & architects Decision‑making frameworks for hardware, parallelism, caching, and cost‑optimization.
Students & interview‑preppers A dedicated section on solving AI system design interview problems with a repeatable 8‑step method.

Core Content Highlights

  1. LLM Inference Mechanics – Prefill vs. decode, TTFT, TPS, KV‑cache, chunked prefill, and disaggregation.
  2. Scaling Strategies – Vertical vs. horizontal scaling, tensor & pipeline parallelism, auto‑scaling warm pools.
  3. Caching Layers – KV‑cache, prompt cache, semantic cache, embedding cache, and compression techniques.
  4. Routing & Load Balancing – Rule‑based, classifier‑based, embedding‑based, LLM‑as‑router, and cascade routing.
  5. RAG & Vector Databases – Index types, hybrid search, HyDE query transformation, reranking, GraphRAG, vector‑less RAG.
  6. Context Management – Window rotation, RoPE decay, summarisation, hierarchical memory.
  7. Agent Architecture – Five core parts, loop engineering, tool calling, Model Context Protocol (MCP), memory stack.
  8. Multi‑Agent Systems – Coordination patterns, A2A protocol, sub‑agents, trade‑offs.
  9. Multimodal & Voice – Edge AI, latency budgeting, barge‑in, streaming via SSE/WebSockets.
  10. Safety & Compliance – Guardrails, prompt injection, red‑team, watermarking, PII redaction, audit logs.
  11. Observability & Evaluation – Traces, spans, LLM‑as‑judge, agent evaluation pipelines.
  12. Cost Optimisation – Model selection, caching, batch inference, cheaper retrieval, self‑hosting.
  13. Interview Framework – 8‑step problem‑solving method, real‑world case studies (Claude Code, Cursor, voice AI), FAQ.

How to Get Started

  1. Clone the repo and open README.md – the table of contents lets you jump to any topic.
  2. Follow the learning path (the Mermaid diagram) if you’re new: start with Foundations → LLM Inference → Scaling → … → Interview.
  3. Deep‑dive blogs are linked from each section; read them for concrete examples, code snippets, and architecture diagrams.
  4. Apply: after each section, try to explain the concept in your own words or prototype a tiny version (e.g., a KV‑cache demo or a simple RAG pipeline).

License

The repository is released under an open‑source license (see the License section at the end of the README), allowing free reuse and modification.


Bottom line: AI System Design is a comprehensive, beginner‑friendly, open‑source guide that teaches engineers how to build, scale, and maintain production AI systems centered on LLMs and agents. It is a genuine educational project in the AI systems space.

Related

  • Project
  • Project
  • Project
  • Project