What we measured Lab-measured
Mean output throughput across four promoted warm p512/n1536, batch-one, 2K-context runs after the strict quality gate. proof file
- Model
- MiniMax M2.7 AutoRound W4A16 INT4
- Publisher
- MiniMax AI
- Checkpoint
- Lasimeri/MiniMax-M2.7-int4-AutoRound
- Compression
- AutoRound W4A16 INT4
- Software
- vLLM XPU + llm-scaler
- Cards
- 4× Intel Arc Pro B70 32 GiB
- Model weight bytes
- —
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding long generation
- Published
- 2026-08-22
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- current clean-host replay
- historical run-time full model payload manifest
- beginner recovery and platform compatibility boundary
- persistent OpenAI-compatible service wrapper for this exact 89 tok/s lane
- decode, prefill, and TTFT context sweep
Measured performance profiles Not published
No qualified structured context or depth profile is published for this package. Diagnostic evidence may still be linked under “What to know” or in the full guide; nothing is estimated in its place — the clearly labeled projection block below is the current best guess.
Many people at once Not published
No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count; the projection below remains clearly labeled as projected.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…