DeepSeek Coder V2 · Lite Instruct

DeepSeek's coding model in its small mixture-of-experts form: 16B parameters total, about 2.4B active per word, built for code completion and programming help. Fits one Arc Pro B70 with room to spare.

57.1tok/s
measured
≈ 43 words a seconddeepseek-coder-v2-lite-instruct · Q4_K_M · llama.cpp SYCL fdb1db877 · 1× B70 · Evidence D of A–D
DeepSeek-Coder-V2 Lite Instruct · base post-trained weights
1 exact artifacts · 1 quantizations
  • deepseek-coder-v2-lite-instruct-8f248fa · Q4_K_M · bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF@8f248fa2072348f77a8bc37754e470de1f61866e
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Runs on 1 card (measured)
Quality evidence
Rapid snapshot only
Measured results
1

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

quant × TP

Exact measured revision and rapid runtime; TP2/4 are unmeasured gaps. Fixed: revision=deepseek-coder-v2-lite-instruct · artifact_id=deepseek-coder-v2-lite-instruct-8f248fa · runtime=llama.cpp SYCL fdb1db877 · mtp=0.

2 untested combinations
TP2·Q4_K_MTP4·Q4_K_M
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Measured speed (quick lab check)

Q4_K_M · one B70 · f16 KV · ctx=2048 configured; this one point is not a context curve

Q4_K_M · tensor parallel cards 157.1tok/sDecode139.83msTTFTstrict rapid snapshot · promoted conservative standalone row · evidence
Q4_K_Mmissing
evidence 1
Fine print

DeepSeek-Coder-V2 Lite sparse MoE. The measured Q4_K_M GGUF is a quantized child artifact of DeepSeek-Coder-V2 Lite Instruct, not a separate model. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? No projection

No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.