MicroLLM Lab Brings 7 Tiny LLMs to the Browser via WebGPU

TL;DR

MicroLLM Lab enables anyone to load, chat with, and benchmark seven small language models (25M–360M parameters) entirely in‑browser via WebGPU, providing sub‑10 ms time‑to‑first‑token latency, zero server cost, and full privacy.

What the Lab Offers

Answer: The site ships a UI that lets users load any of seven Q4‑quantized models, run inference on‑device, and view benchmark results without any backend.

  • Models: PetitGPT‑research‑v1 (125 M), SmolLM2 135M Instruct, L20‑Edu 135M, SmolLM2 360M Instruct, MiniMind2 104M, MiniMind2 Small 26M, GPT‑2 124M (nanoGPT‑shaped). All models are Apache‑2.0 or MIT licensed and range from 15 MB to 216 MB after 4‑bit quantization.
  • Hardware acceleration: When WebGPU is available, inference runs on the GPU at 100–300 tokens / s (10–20× faster than the WASM fallback’s 8–20 tok/s). The UI reports the current speed and can generate a shareable performance certificate.
  • Privacy & cost: All computation happens in the browser; prompts never leave the device, and there are no API fees.
  • Use‑case focus: Edge classification, spam filtering, intent extraction, and any task where a cheap, ultra‑low‑latency model can triage before calling a large cloud LLM.

How to Get Started

Answer: Loading a model is a three‑step process.

  1. Load – Click the Load button on a model card; the model is cached in IndexedDB (no file‑system download).
  2. Chat – Select the loaded model and use the chat panel to issue prompts; token‑per‑second speed is displayed live.
  3. Benchmark – Switch to the Benchmarks tab to run a suite of objective tests (regex‑based checks) and generate a performance certificate that records peak and sustained token rates.

The UI also provides a Custom eval editor where users can write JavaScript benchmark definitions.

Technical Foundations

Answer: The Lab relies on four key technologies.

  • Small Language Model (SLM): A compact transformer (25M–360M parameters) designed for edge inference rather than broad knowledge.
  • Q4 Quantization: Weights are stored in 4‑bit integers, shrinking a 100 M‑parameter model to ~50 MB while preserving generation quality.
  • WebGPU: A W3C standard that maps compute shaders to native GPU APIs (Metal, DirectX 12, Vulkan). When available, it delivers 10–20× speedups over the WASM fallback.
  • IndexedDB Cache: Models persist across sessions in the browser’s private storage, eliminating repeated downloads.

Performance Observations from Users

Answer: Community feedback highlights both strengths and pain points.

"Impressive how much you can do locally now. Was expecting much slower inference but it's surprisingly snappy." – Mbarley

"The speed for in‑browser execution is wild. Perfect for quick demos without a backend." – hbroom

"Running on CPU fallback yields only 8–20 tok/s; with WebGPU it jumps to 100–300 tok/s, a 10–20× boost." – Lab documentation

Issues reported:

  • UI density and small text make the interface hard to navigate (comment by demibabs).
  • Some browsers (Firefox) lack WebGPU support, leading to errors like GPUShaderStage is not defined (comment by anonimous-emacs and langurmonkey).
  • Model outputs can be nonsensical or repetitive, especially on larger prompts (comments by not2b, Schlagbohrer, botanrice).

Sample Model Behaviors

Answer: The models demonstrate varied competence.

  • PetitGPT research‑v1 correctly solves arithmetic but sometimes adds extra steps: "2 + 2 = 4 … So, 2 + 2 = 4 + 2."
  • SmolLM2 360M provides wildly inaccurate facts (e.g., claiming California’s population is 153 million).
  • L20‑Edu can generate plausible‑looking code snippets but may repeat filler text.

These examples illustrate that while latency is excellent, factual accuracy remains limited.

Community Extensions & Related Projects

Answer: Several developers are building on the same idea of in‑browser LLMs.

  • tiny.tobelabs.com – A minimal UI for a Tiny Stories model and a larger QA model (comment by atobe).
  • sonistellar.com/lab – Another WebGPU‑based LLM playground (comment by vs4vijay).
  • Web Models API – A proposal for a standardized browser API to run open‑weight models on‑device (comment by kenzic).
  • Three‑LLM – Uses ThreeJS’s shading language to run LLMs in the browser (comment by bhouston).

Practical Takeaways

Answer: MicroLLM Lab showcases that 4‑bit‑quantized SLMs can run at interactive speeds on modern browsers, making edge AI feasible for low‑cost, privacy‑preserving applications.

  • When to use: Real‑time classification, intent routing, or prototyping where latency and cost matter more than deep factual knowledge.
  • When not to use: Tasks requiring high factual accuracy, multi‑language support, or complex reasoning; the models often hallucinate or repeat.
  • Future direction: Wider WebGPU adoption (especially in Firefox) and improved UI/UX will broaden accessibility. Standardizing on‑device model APIs could further lower the barrier for developers.

All information is derived from the MicroLLM Lab site and the Hacker News discussion thread linked above.

Sources

Related

  • Dispatch
  • Project
  • Dispatch
  • Project
  • Project