Packets and recipes
The deployment variants of this family, at every maturity.
MiniMax M2.7 AutoRound · 4 cards, fully checked
Reproduce 89.31 tok/s on 4× B70 · historical-strict-speed-candidate
Mean output throughput across four promoted warm p512/n1536, batch-one, 2K-context runs after the strict quality gate.MiniMax M2.7 AutoRound · deployable TP4
Evidence packet · qualified-deployment-reference
MiniMax M2.7 · production c1 profile
Evidence packet · banked-service-profile
lab report — documents the result; not a step-by-step install guideMiniMax M2.7 · constrained HTML
Evidence packet · research-only-task-specific
lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
AutoRound · target only
MiniMax M2.7 AutoRound W4A16 · vLLM XPU · graph on · f16-family KV · MTP/speculation off. Only TP4 is measured in the maintained family packet.
- ✓ Measured4 cards, no speculative decoding
TP4·MTP089.53 tok/sevidence
2 untested combinations
TP1·MTP0TP2·MTP0Measured results
Every number links to its proof.
Deployment options
TP4 · p512/n1536 comparison gate · active context 2K; x is configured service capacity, not active prompt length
Speed at a 32K-token conversation
TP4 · configured 32K · warm p510/n1536 request; prefill is a conservative lower bound
Speed in real service
TP4 production c1 · different output lengths at each prompt size; discrete capacity observations, not an interpolated context curve
Prompt processing speed
TP4 OpenAI endpoint · max_tokens=1 · includes HTTP and first-token work, so values are conservative lower bounds
Fine print
MiniMaxM2ForCausalLM, 62 layers, 3072 hidden. one packaged AutoRound revision; workload and runtime identities remain separate. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
2 combinations the lab has stopped pursuing: The 33,792-token startup loaded and entered compile/warmup but did not expose /v1/models within the wait window; 32,768 is the promoted service setting. evidence The larger prefill chunk triggered the documented PIECEWISE compile-range/startup failure; 512 remains the stable service setting. evidence Exact selectors are in the family data.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (minimax-m2.7-lasimeri · AutoRound W4A16 INT4 · vLLM XPU 0.20.1-local · 4× B70). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…