DeepSeek V4 · Flash 180B

DeepSeek's large mixture-of-experts reasoning model in a community-trimmed 180B form (about 13B active per word). A research-grade deployment: it needs all four B70 cards and an experimental checkpoint, so treat it as a look at what the biggest open models can do on this hardware, not a daily driver.

78.29tok/s
measured
≈ 59 words a seconddeepseek-v4-flash-180b-community · experimental uniform-K160 FP8/MXFP4 target · vLLM XPU record stack 264c7f2/3131567/48fda4f · 4× B70 · Evidence C of A–D
DeepSeek V4 Flash 180B · community base identity · parent model identity
1 exact artifacts · 1 quantizations
  • deepseek-v4-flash-180b-k160-7c360e1 · experimental uniform-K160 FP8/MXFP4 target · 0xSero/DeepSeek-V4-Flash-180B@7c360e1cd4a5168099dbc54d16d929bf6df04990
B70 fit
Runs on 4 cards (measured)
Quality evidence
Target verified exact gates
Measured results
3

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

attached MTP depth × TP

Exact experimental K160 artifact · PIECEWISE · FP8 KV. Historical TP4 rows use EP4 and retain their exact differing runtime identities; rates are not transferred across them. The approximately 103.1 GB on-disk target does not fit fully resident at TP1 or TP2 on 32/64 GiB; CPU offload is outside this view. Fixed: revision=deepseek-v4-flash-180b-community · artifact_id=deepseek-v4-flash-180b-k160-7c360e1 · variant=experimental uniform-K160 FP8/MXFP4 target · quantization=experimental uniform-K160 FP8/MXFP4 target · graph_mode=PIECEWISE · kv=fp8.

  • ✓ Measured4 cards, target-only TP4·target-only43.77 tok/sevidence
  • ✓ Measured4 cards, MTP1 TP4·MTP163.85 tok/sevidence
  • ■ Closed4 cards, MTP2 TP4·MTP2 — M=3 capture and 10/10 exact pre-hang rows passed, but the second draft accepted only 0.5–2.2% and realistic traffic stopped before a valid suite completed.no valid rate
  • ■ Closed4 cards, MTP4 TP4·MTP4 — The evidence-based post-MTP3 stop decision closed four drafted tokens after the third proposal accepted at most 3.2%. No MTP4 throughput exists.not launched
  • × Unsupported1 card, target-only TP1·target-only — The approximately 103.1 GB on-disk K160 target cannot fit fully resident on one 32 GiB B70; CPU offload is outside this view.weights do not fit
  • × Unsupported2 cards, target-only TP2·target-only — The approximately 103.1 GB on-disk K160 target cannot fit fully resident on two 32 GiB B70s; CPU offload is outside this view.weights do not fit
  • × Unsupported1 card, MTP1 TP1·MTP1 — The fully resident target already exceeds one B70 before the attached MTP layer is considered.weights do not fit
  • × Unsupported2 cards, MTP1 TP2·MTP1 — The fully resident target already exceeds two B70s before the attached MTP layer is considered.weights do not fit
  • × Unsupported1 card, MTP2 TP1·MTP2 — The fully resident target already exceeds one B70; repeated-layer drafting cannot change target capacity.weights do not fit
  • × Unsupported2 cards, MTP2 TP2·MTP2 — The fully resident target already exceeds two B70s; repeated-layer drafting cannot change target capacity.weights do not fit
  • × Unsupported1 card, MTP3 TP1·MTP3 — The fully resident target already exceeds one B70; repeated-layer drafting cannot change target capacity.weights do not fit
  • × Unsupported2 cards, MTP3 TP2·MTP3 — The fully resident target already exceeds two B70s; repeated-layer drafting cannot change target capacity.weights do not fit
  • × Unsupported1 card, MTP4 TP1·MTP4 — The fully resident target already exceeds one B70; repeated-layer drafting cannot change target capacity.weights do not fit
  • × Unsupported2 cards, MTP4 TP2·MTP4 — The fully resident target already exceeds two B70s; repeated-layer drafting cannot change target capacity.weights do not fit
  • ⚠ Quarantined4 cards, MTP3 TP4·MTP3 — Only 8/10 rows reached the timing window; the third draft accepted 0–3.2%, so the 46.247281 diagnostic is not a valid endpoint result. Observed: D46.247 · incomplete.evidence

Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.

speculation × TP

Exact experimental K160 target and record runtime; missing cells have no inherited result. Fixed: revision=deepseek-v4-flash-180b-community · artifact_id=deepseek-v4-flash-180b-k160-7c360e1 · variant=experimental uniform-K160 FP8/MXFP4 target · runtime=vLLM XPU record stack 264c7f2/3131567/48fda4f · graph_mode=PIECEWISE.

  • ✓ Measured4 cards, DSpark7 TP4·DSpark780.82 tok/sguide · evidence
  • × Unsupported1 card, DSpark7 TP1·DSpark7 — The approximately 103.1 GB on-disk K160 target cannot fit fully resident on one 32 GiB B70, before the DSpark draft pack is considered; CPU offload is outside this view.weights do not fit
  • × Unsupported2 cards, DSpark7 TP2·DSpark7 — The approximately 103.1 GB on-disk K160 target cannot fit fully resident on two 32 GiB B70s, before the DSpark draft pack is considered; CPU offload is outside this view.weights do not fit

Measured results

Every number links to its proof.

family data

Measured speed on four cards

Decode retains three independent suite medians (76.900178–80.820052); the packet headline is the 78.287226 median-of-medians, while TTFT is the 342.777657 record-suite observation

experimental K160 + DSpark7 · tensor parallel cards 476.9–80.82tok/sDecode342.78msTTFTthree independent strict suites · closed frontier record · evidence
experimental K160 + DSpark7missing
evidence 1

DSpark width screen ◇ screened, experimental

Same patched DEV binary and K160 target · TP4+EP4 · PIECEWISE · FP8 KV · public+DEV prompts · screened, not promotion evidence

022.344.76789.45678DSpark draft tokens · tok/sDSpark DEV widths · Decode (decode_tok_s) · DSpark draft tokens=5.0 · value=76.314666 tok/sDSpark DEV widths · Decode (decode_tok_s) · DSpark draft tokens=6.0 · value=72.803326 tok/sDSpark DEV widths · Decode (decode_tok_s) · DSpark draft tokens=7.0 · value=79.801765 tok/sDSpark DEV widths · Decode (decode_tok_s) · DSpark draft tokens=8.0 · value=73.834972 tok/s
DSpark DEV widths
DSpark DEV widths 5: 76.31 · 6: 72.8 · 7: 79.8 · 8: 73.83
evidence 1
Fine print

DeepSeek V4 Flash 180B sparse MoE. The experimental K160 export is a quantized child artifact of the DeepSeek V4 Flash community model identity, not a separate model. Measurements remain artifact- and runtime-specific. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

4 combinations the lab has stopped pursuing: Repeated use of the attached single MTP layer stopped before a valid suite completed; its second proposal accepted only 0.5–2.2% on realistic traffic. evidence The three-draft/four-row endpoint produced only eight eligible timing rows at 46.247281 tok/s; the third proposal accepted 0–3.2%, so the incomplete screen is diagnostic only. evidence True MTP4 was not launched because the evidence-based post-MTP3 decision rejected deeper repeated-layer work after near-zero third-proposal acceptance. evidence The configuration frontier is paused; reopen only for a substantial EAGLE capture/training effort or a genuinely new mechanism, not another generic flag sweep. evidence Exact selectors are in the family data.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (deepseek-v4-flash-180b-community · experimental uniform-K160 FP8/MXFP4 target · vLLM XPU record stack · 4× B70). The grade is optimization headroom against the tuned-run target, not model quality.