MiniMax M2.7

MiniMax M2.7, a 229B mixture-of-experts (about 10B active per word) built for long conversations and agent-style work. Needs all four Arc Pro B70 cards in 4-bit form.

89.31tok/s
measured
≈ 67 words a secondminimax-m2.7-lasimeri · AutoRound W4A16 INT4 · vLLM XPU 0.20.1-local · 4× B70 · B70-verified historical
MiniMax M2.7 · Lasimeri AutoRound export · measured deployment weights
B70 fit
Four card
Quality evidence
Strong scoped evidence
Measured results
6

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

AutoRound · target only

MiniMax M2.7 AutoRound W4A16 · vLLM XPU · graph on · f16-family KV · MTP/speculation off. Only TP4 is measured in the maintained family packet.

  • ✓ Measured4 cards, no speculative decoding TP4·MTP089.53 tok/sevidence
2 untested combinations
TP1·MTP0TP2·MTP0
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Deployment options

TP4 · p512/n1536 comparison gate · active context 2K; x is configured service capacity, not active prompt length

historical strict stack · configured max context tokens 2K88.93–89.53tok/sDecodelab measurement; quality scope in evidence · evidence
deployable 32K stack · configured max context tokens 32K82.37–83.55tok/sDecodelab measurement; quality scope in evidence · evidence
historical strict stackdeployable 32K stack
evidence 1 · evidence 2

Speed at a 32K-token conversation

TP4 · configured 32K · warm p510/n1536 request; prefill is a conservative lower bound

OpenAI endpoint · active prompt tokens 51085.45tok/sDecode1,446tok/sPrefill351.07msTTFTlab measurement; quality scope in evidence · evidence
OpenAI endpoint
evidence 1

Speed in real service

TP4 production c1 · different output lengths at each prompt size; discrete capacity observations, not an interpolated context curve

027.154.281.3108.437218.9K32.3Kprompt tokens · tok/sproduction c1 · Decode (decode_tok_s) · prompt tokens=372.0 · value=96.82 tok/sproduction c1 · Decode (decode_tok_s) · prompt tokens=18924.0 · value=77.89 tok/sproduction c1 · Decode (decode_tok_s) · prompt tokens=32264.0 · value=63.91 tok/s
production c1
production c1 372: 96.82 · 18.9K: 77.89 · 32.3K: 63.91
evidence 1

Prompt processing speed

TP4 OpenAI endpoint · max_tokens=1 · includes HTTP and first-token work, so values are conservative lower bounds

0512.51,0251,5372,0505102K8.1K16.3Kprompt tokens · tok/swarm lower bound · Prefill (prefill_tok_s) · prompt tokens=510.0 · value=1607.9 tok/swarm lower bound · Prefill (prefill_tok_s) · prompt tokens=2036.0 · value=1750.1 tok/swarm lower bound · Prefill (prefill_tok_s) · prompt tokens=8144.0 · value=1830.3 tok/swarm lower bound · Prefill (prefill_tok_s) · prompt tokens=16288.0 · value=1810.3 tok/s
warm lower bound
warm lower bound 510: 1,608 · 2K: 1,750 · 8.1K: 1,830 · 16.3K: 1,810
evidence 1
Fine print

MiniMaxM2ForCausalLM, 62 layers, 3072 hidden. one packaged AutoRound revision; workload and runtime identities remain separate. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

2 combinations the lab has stopped pursuing: The 33,792-token startup loaded and entered compile/warmup but did not expose /v1/models within the wait window; 32,768 is the promoted service setting. evidence The larger prefill chunk triggered the documented PIECEWISE compile-range/startup failure; 512 remains the stable service setting. evidence Exact selectors are in the family data.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (minimax-m2.7-lasimeri · AutoRound W4A16 INT4 · vLLM XPU 0.20.1-local · 4× B70). The grade is optimization headroom against the tuned-run target, not model quality.