Candidate package · vLLM XPU + llm-scaler · 4× Intel Arc Pro B70

MiniMax M2.7 AutoRound INT4 on four Intel Arc Pro B70 cards

MiniMax M2.7, a 229B mixture-of-experts for long conversations, on four Arc Pro B70 cards in 4-bit form with vLLM. Quality-checked token by token.

What we measured Lab-measured

89.3tok/s mean output throughput

Mean output throughput across four promoted warm p512/n1536, batch-one, 2K-context runs after the strict quality gate. proof file

Model
MiniMax M2.7 AutoRound W4A16 INT4
Publisher
MiniMax AI
Checkpoint
Lasimeri/MiniMax-M2.7-int4-AutoRound
Compression
AutoRound W4A16 INT4
Software
vLLM XPU + llm-scaler
Cards
4× Intel Arc Pro B70 32 GiB
Model weight bytes
Operating systems
Linux
Delivery
native
Good for
general coding long generation
Published
2026-08-22
Clean-host replay
not yet

Still missing before this becomes an install guide

  • current clean-host replay
  • historical run-time full model payload manifest
  • beginner recovery and platform compatibility boundary
  • persistent OpenAI-compatible service wrapper for this exact 89 tok/s lane
  • decode, prefill, and TTFT context sweep

Measured performance profiles Not published

No qualified structured context or depth profile is published for this package. Diagnostic evidence may still be linked under “What to know” or in the full guide; nothing is estimated in its place — the clearly labeled projection block below is the current best guess.

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count; the projection below remains clearly labeled as projected.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.