DwarfStar 4 (ds4) Enables Local Inference of Frontier LLMs on High‑Memory Macs and GPUs
TL;DR – What ds4 delivers and why it matters
DwarfStar 4 (ds4) lets you run frontier‑weight large language models such as DeepSeek V4/V4.1, GLM 5.x and Qwen 3.8 locally on high‑memory Macs, NVIDIA CUDA, or AMD ROCm hardware by compressing MoE experts with asymmetric 2‑bit quantization and streaming KV‑cache to SSD. This makes inference with 284‑billion‑parameter models feasible on a single workstation, removing the need for costly remote API calls.
Core design – a purpose‑built inference stack
ds4 is a narrow, C‑only engine that targets a small, well‑defined set of MoE models and validates each layout end‑to‑end. Unlike generic GGUF runners, ds4 focuses on three model families (DeepSeek V4/V4.1, GLM 5.x, Qwen 3.8) and provides three unified interfaces: a CLI (./ds4), a local HTTP server (./ds4‑server), and a persistent coding agent (./ds4‑agent).
Asymmetric 2‑bit quantization
"Compress the routed experts, keep critical shared paths precise. That is how the supported routed‑MoE builds fit their target machines." – ds4 documentation
- Experts in the mixture‑of‑experts (MoE) layers are quantized to 2 bits.
- Shared routing and attention pathways retain higher precision, preserving model quality.
- The technique enables a 284 B‑parameter DeepSeek V4 model to run on a 128 GB system.
KV‑cache as a disk citizen
"Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re‑prefill." – ds4 documentation
- Prefix KV‑cache can be flushed to SSD, allowing contexts up to 65 k tokens without exhausting RAM.
- Prompt‑hash‑based resumption avoids re‑prefilling after a restart, crucial for long‑running agents.
Unified three‑interface model state
./ds4– interactive chat CLI../ds4‑server– OpenAI‑compatible HTTP API for editors, agents, or custom clients../ds4‑agent– persistent coding session with built‑in state management.
Supported models and hardware fit
ds4 currently supports:
- DeepSeek V4 & V4.1 Flash (Mixture‑of‑Experts, 284 B parameters)
- GLM 5.x (including Flash variants)
- Qwen 3.8 Flash Next (vision‑enabled optional)
Memory requirements and performance
| Platform | Memory | Model (quant) | Context | Prefill t/s | Generation t/s |
|---|---|---|---|---|---|
| Apple M5 Max | 128 GB | DeepSeek V4 Q2 | 2 048 tok | 790.2 | 39.4 |
| Apple M5 Max | 128 GB | DeepSeek V4 Q2 | 65 536 tok | 398.5 | 27.6 |
| NVIDIA DGX Spark | 128 GB | DeepSeek V4 Q2 | 2 048 tok | 825.8 | 18.1 |
| NVIDIA DGX Spark | 128 GB | DeepSeek V4 Q2 | 65 536 tok | 823.0 | 13.8 |
The benchmark table is taken directly from the ds4 website.
Hardware matrix – The baseline model (DeepSeek V4 Q2) runs comfortably on any 128 GB system. GLM 5.3 Q2 and Qwen Q4 also fit at 128 GB, while DeepSeek V4.1 Q2 streams from SSD, allowing operation on 64 GB Macs.
Quick‑start workflow
- Fetch the GGUF weights
git clone https://github.com/antirez/ds4 cd ds4 && ./download_model.sh ds4f-q2 - Build for your backend (Metal, CUDA, or ROCm)
make # generic build (Metal on macOS) make cuda-spark # CUDA‑optimized build - Run
./ds4 # interactive CLI ./ds4-server --ctx 100000 # start OpenAI‑compatible API
ds4 does not aim to run arbitrary GGUF files; only the verified model layouts are supported.
Community extensions and ecosystem
- ds4go – a Go FFI wrapper and TUI that lets other languages call ds4 as a shared library. Maintained by @neomantra, it adds Vision and Qwen support and provides Homebrew‑installable binaries. (GitHub release v0.8.20260)
- Club‑3090 server – a web frontend for ds4 maintained by @gchamon, enabling remote access to the local engine.
- Custom SSD‑cache forks – users have back‑ported the disk‑cache feature to llama.cpp and other runners (e.g., @xlayn’s fork).
How ds4 differs from other local runners
| Feature | ds4 | llama.cpp / ollama |
|---|---|---|
| Targeted MoE models (DeepSeek V4, GLM 5.x, Qwen 3.8) | ✅ | ❌ (generic GGUF only) |
| Asymmetric 2‑bit quantization of experts | ✅ | ❌ (usually uniform quant) |
| KV‑cache persisted to SSD | ✅ | ❌ (in‑memory only) |
| Unified CLI, server, and agent sharing state | ✅ | Partial (separate binaries) |
| Metal‑first implementation for Apple Silicon | ✅ | Limited (mostly CPU) |
Community members note that ds4’s minimal‑dependency philosophy mirrors Redis, making it lightweight and easy to embed.
Open questions and community feedback
- Tool‑calling performance – Users such as @cuttothechase ask for TPS numbers when using function calling; current benchmarks focus on raw pre‑fill and generation speed.
- Model quality after quantization – @doctorpangloss mentions concerns about the DeepSeek V4 checkpoint quality after aggressive quantization.
- Minimum RAM on Apple Silicon – Discussions reveal some confusion: the website cites 64 GB as a lower bound for SSD‑streamed runs, while the GitHub README mentions 96 GB for full‑speed Metal execution.
- Comparison to other runners – Several commenters (e.g., @locknitpicker) request a side‑by‑side speed comparison with ollama or llama.cpp; ds4’s niche is MoE‑specific optimizations rather than universal GGUF support.
Verdict
ds4 provides a practical pathway to run frontier‑weight MoE models locally on a single workstation, leveraging asymmetric quantization and SSD‑backed KV‑caching to overcome memory limits. Its focused model support, low‑dependency C implementation, and three unified interfaces make it a compelling alternative to generic runners for users who need high‑quality inference without cloud latency.
Sources
Related
- Dispatch
- Project
- Dispatch
- Dispatch
- Project