ds4 (DwarfStar 4), a small MIT licensed C inference engine, runs frontier mixture of experts AI models (architectures that activate only a small slice of weights per token) on a high memory Mac.
ds4 is a narrow MIT-licensed C inference engine that applies 2-bit compression to only the routed-expert weights of frontier mixture-of-experts models, keeping shared paths precise. That specific lever is what puts DeepSeek V4 Flash Q2 onto a 64GB Mac and GLM 5.3 Q2 plus Qwen Q4 onto 128GB. Three interfaces share one model state: a ./ds4 CLI, a ./ds4-server for local HTTP APIs, and a ./ds4-agent for persistent coding sessions.
The project is opinionated on purpose. Where tools like llama.cpp and Ollama chase broad model coverage, ds4 validates a small set of routed-MoE builds end to end: DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen 3.8 Flash Next, with text and vision on select families. The trade is coverage for fit. Asymmetric quantization is the technical reason: experts are sparse-activated during inference, so most of them sit idle on any given token, and turning their weights down to 2 bits costs little quality because they are rarely the bottleneck. The architecture docs describe that split explicitly.
The memory tier ladder the engine targets is concrete. V4 Flash Q2 is the baseline fit at 64GB; 128GB opens up GLM 5.3 Q2 and Qwen Q4; V4.1 Q2 streams from SSD, which is the project's escape valve for long-context workloads. SSD prefix streaming is keyed by prompt hash so a restart can skip the full prefill step on a cache hit. The engine layers tensor parallelism, session batching, DSpark, and MTP on top of that.
Self-reported performance is the obvious place to slow down. The benchmark table lists a M5 Max 128GB at 32K context, V4 Flash Q2, at 34.4 tokens per second generation and 557 tokens per second prefill. Those are author numbers, not independent reproductions, and the narrow model support is part of the same bet: validating each layout end to end is what makes those rows checkable at all. Independent third-party work is already in motion. The Hacker News discussion includes an outside maintainer who has exposed ds4 as shared libraries with public builds, and a Go FFI binding called ds4go, plus community-reported additions of Vision and Qwen support.
The framing the source carries is that this comes from Salvatore Sanfilippo, the creator of Redis, an attribution consistent with the Hacker News thread but not restated on the captured project page and not yet independently verified against a signed commit or repository. Treat it as a hook, not a foundation. The lever is what survives the verification step: a specific, named architectural choice that converts a class of frontier weights from "needs a GPU rack" to "fits on a high-memory laptop."
Two things to watch. The named model list must check out: DeepSeek V4, V4.1 Flash, GLM 5.x, and Qwen 3.8 Flash Next are 2026-10-window branding, and "frontier on a laptop" only lands if those model cards actually exist as advertised. The mechanism is sound; the model list is the load-bearing claim.
The interesting part is the engineering choice, not the brand: pick a narrow class of frontier models, find the specific weight class that tolerates aggressive compression, validate the layout end to end, and ship. If the asymmetric routed-expert bet holds up under independent reproduction, local frontier inference has a concrete mechanism behind it.