DS4 · LOCAL FRONTIER INFERENCE
DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.
SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB
ds4 · local session
PRINCIPLE · LOCAL MODEL STACK
PHASE 1 · THE GIANT
DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.
PHASE 2 · THE COLLAPSE
Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.
PHASE 3 · THE DWARF STAR
The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.
SCROLL ▾
DESIGN CHOICES
Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.
CORE 01
Compress the routed experts, keep critical shared paths precise. That is how the supported routed-MoE builds fit their target machines.
CORE 02
Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re-prefill.
CORE 03
Use ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions.
ARCHITECTURE
Project GGUFs, a self-contained engine and agent-facing interfaces, checked against official model outputs.
MODEL DeepSeek V4 / V4.1 GLM 5.x · Qwen3.8 supported GGUF layouts only asymmetric 2-bit + imatrix load ENGINE ds4 engine written in C metal · cuda · rocm parallel · batch · speculate KV cache RAM ⇄ SSD · survives restarts serve ./ds4 interactive CLI ./ds4-server OpenAI + Anthropic API ./ds4-agent native coding agent personal → distributed memory classes
RUNTIME MAP · SIMPLIFIED. SEE ARCHITECTURE NOTES FOR THE FULL DRAWING.
RUN IT
Download the project GGUF, build for your backend, then start the CLI or server. Generic GGUF files are not the target.
STEP 1 · FETCH THE WEIGHTS
ds4 · zsh
$ git clone https://github.com/antirez/ds4
$ cd ds4 && ./download_model.sh ds4f-q2
STEP 2 · BUILD FOR YOUR BACKEND
ds4 · zsh
$ make
$ make cuda-spark
STEP 3 · TALK TO IT
ds4 · zsh
$ ./ds4
$ ./ds4-server --ctx 100000
FIT CHECK
Pick your platform and memory: get a conservative starting path and understand which execution modes apply.
PLATFORM MEMORY
✓ Runs well
V4 Flash Q2 is the baseline. At 128 GB, GLM 5.3 Q2 and Qwen Q4 also fit; V4.1 Q2 streams from SSD.
./download_model.sh ds4f-q2 && make
REF · M5 MAX 128GB · 32K CTX: 34.4 T/S GEN · 557 T/S PREFILL
Estimates from the ds4 benchmark table. Full guide in Hardware and Installation.
BENCHMARKS
Reference rows from upstream. Read prefill and generation separately, especially for long-context agent workloads.
| Machine | Context | Prefill t/s | Generation t/s |
|---|---|---|---|
| M5 Max, 128 GB | q2 · 2,048 tok | 790.2 | 39.4 |
| M5 Max, 128 GB | q2 · 65,536 tok | 398.5 | 27.6 |
| DGX Spark, 128 GB | q2 · 2,048 tok | 825.8 | 18.1 |
| DGX Spark, 128 GB | q2 · 65,536 tok | 823.0 | 13.8 |
API & AGENTS
ds4-server speaks OpenAI and Anthropic-style APIs, so local coding agents can connect to your own machine with a base URL.
Start with the quickstart, check the hardware matrix, then connect your editor, agent or API client to the local server.