amitshekhariitbhu/ai-system-design
AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step.
AI System Design – What This Repo Is
Type: Open‑source learning guide (markdown‑heavy repository) Scope: Covers the whole AI system design discipline – from LLM inference mechanics, hardware choices, caching, routing, and RAG, to agents, multi‑agent orchestration, multimodal/voice pipelines, safety guardrails, observability, cost optimisation and interview preparation.
Why it’s in‑scope: The repository is a curated, continuously‑updated collection of explanatory articles, diagrams and links that teach how to build production‑grade AI services. All topics are directly related to modern AI/LLM engineering, so it qualifies as a genuine AI‑focused educational project.
Quick Summary
- A step‑by‑step guide for engineers who want to design, deploy, and scale AI systems built around large language models.
- Organized as a long Markdown document with a detailed table of contents that links to deep‑dive blog posts for each sub‑topic.
- Covers foundations (tokens, GPUs, inference servers) → core engineering (scaling, caching, routing, RAG) → advanced layers (agents, multi‑agent systems, multimodal, safety, observability, cost, fine‑tuning) → interview framework.
- Written for beginners, avoids jargon, and includes real‑world numbers, tools, and architecture diagrams.
- Maintained by Amit Shekhar, founder of Outcome School, with cross‑links to his teaching platform and related AI engineering course.
Who Might Use This
| Audience | What They Get |
|---|---|
| Software engineers moving into AI | A clear roadmap from basic token economics to production‑grade pipelines. |
| Backend / mobile / frontend developers | Practical guidance on integrating LLM inference, streaming, rate‑limiting, and gateways. |
| ML / data scientists | Steps to transition models from research to a scalable serving stack (GPU sizing, quantisation, etc.). |
| Engineering managers & architects | Decision‑making frameworks for hardware, parallelism, caching, and cost‑optimization. |
| Students & interview‑preppers | A dedicated section on solving AI system design interview problems with a repeatable 8‑step method. |
Core Content Highlights
- LLM Inference Mechanics – Prefill vs. decode, TTFT, TPS, KV‑cache, chunked prefill, and disaggregation.
- Scaling Strategies – Vertical vs. horizontal scaling, tensor & pipeline parallelism, auto‑scaling warm pools.
- Caching Layers – KV‑cache, prompt cache, semantic cache, embedding cache, and compression techniques.
- Routing & Load Balancing – Rule‑based, classifier‑based, embedding‑based, LLM‑as‑router, and cascade routing.
- RAG & Vector Databases – Index types, hybrid search, HyDE query transformation, reranking, GraphRAG, vector‑less RAG.
- Context Management – Window rotation, RoPE decay, summarisation, hierarchical memory.
- Agent Architecture – Five core parts, loop engineering, tool calling, Model Context Protocol (MCP), memory stack.
- Multi‑Agent Systems – Coordination patterns, A2A protocol, sub‑agents, trade‑offs.
- Multimodal & Voice – Edge AI, latency budgeting, barge‑in, streaming via SSE/WebSockets.
- Safety & Compliance – Guardrails, prompt injection, red‑team, watermarking, PII redaction, audit logs.
- Observability & Evaluation – Traces, spans, LLM‑as‑judge, agent evaluation pipelines.
- Cost Optimisation – Model selection, caching, batch inference, cheaper retrieval, self‑hosting.
- Interview Framework – 8‑step problem‑solving method, real‑world case studies (Claude Code, Cursor, voice AI), FAQ.
How to Get Started
- Clone the repo and open
README.md– the table of contents lets you jump to any topic. - Follow the learning path (the Mermaid diagram) if you’re new: start with Foundations → LLM Inference → Scaling → … → Interview.
- Deep‑dive blogs are linked from each section; read them for concrete examples, code snippets, and architecture diagrams.
- Apply: after each section, try to explain the concept in your own words or prototype a tiny version (e.g., a KV‑cache demo or a simple RAG pipeline).
License
The repository is released under an open‑source license (see the License section at the end of the README), allowing free reuse and modification.
Bottom line: AI System Design is a comprehensive, beginner‑friendly, open‑source guide that teaches engineers how to build, scale, and maintain production AI systems centered on LLMs and agents. It is a genuine educational project in the AI systems space.
Related
- Project
- Project
- Project
- Project