docs(minimax-m3): refresh B200 benchmarks (tp8, piecewise) + add GPQA (#28207)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
This commit is contained in:
@@ -42,7 +42,7 @@ docker pull lmsysorg/sglang:dev-cu13-minimax-m3
|
|||||||
The command panel below fills in the right tag per platform: `dev-cu13-minimax-m3` (CUDA 13 — B300, GB200, GB300), `dev-cu12-minimax-m3` (CUDA 12 — Hopper H200), or `dev-minimax-m3` (default). On AMD Instinct it uses the matching ROCm image (MI300X/MI325X → `…-rocm700-mi30x`, MI350X/MI355X → `…-rocm720-mi35x`). For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker), substituting the inner `sglang serve ...` with what the command generator produces.
|
The command panel below fills in the right tag per platform: `dev-cu13-minimax-m3` (CUDA 13 — B300, GB200, GB300), `dev-cu12-minimax-m3` (CUDA 12 — Hopper H200), or `dev-minimax-m3` (default). On AMD Instinct it uses the matching ROCm image (MI300X/MI325X → `…-rocm700-mi30x`, MI350X/MI355X → `…-rocm720-mi35x`). For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker), substituting the inner `sglang serve ...` with what the command generator produces.
|
||||||
|
|
||||||
<Note>
|
<Note>
|
||||||
These images do not yet bundle MiniMax's MSA sparse-attention kernel (bundling it by default is in progress). Blackwell users who want the recommended fast path install MSA manually after pulling — see **§2.1**. Without it, the same recipe still serves on the built-in Triton sparse path.
|
These M3 dev images now **bundle MiniMax's MSA sparse-attention kernel** (`fmha_sm100`), so Blackwell users get the recommended fast path automatically — no manual install needed (see **§2.1**). On a custom image without it, the same recipe still serves on the built-in Triton sparse path.
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
</Tab>
|
</Tab>
|
||||||
@@ -79,7 +79,7 @@ Key characteristics as served by SGLang:
|
|||||||
- **Sparse attention**: most layers use M3's "lightning indexer" block-sparse attention (top-k 128-token blocks), which keeps decode cost roughly flat in context length. On Blackwell, MiniMax's open-source [MSA kernel](https://github.com/MiniMax-AI/MSA) accelerates this path further (§2.1).
|
- **Sparse attention**: most layers use M3's "lightning indexer" block-sparse attention (top-k 128-token blocks), which keeps decode cost roughly flat in context length. On Blackwell, MiniMax's open-source [MSA kernel](https://github.com/MiniMax-AI/MSA) accelerates this path further (§2.1).
|
||||||
- **MXFP8 quantization across vendors**: the MXFP8 MoE weights run natively on NVIDIA Blackwell (B200 / B300 / GB200 / GB300) and on AMD Instinct MI350X/MI355X (gfx950 / CDNA4), both of which have hardware MX-scaled matmul. On AMD MI300X/MI325X (gfx942 / CDNA3) — no hardware MX — SGLang converts the weights to block-fp8 `[128,128]` at load and serves them on the tuned ROCm kernels (§2.3). The vision tower stays unquantized.
|
- **MXFP8 quantization across vendors**: the MXFP8 MoE weights run natively on NVIDIA Blackwell (B200 / B300 / GB200 / GB300) and on AMD Instinct MI350X/MI355X (gfx950 / CDNA4), both of which have hardware MX-scaled matmul. On AMD MI300X/MI325X (gfx942 / CDNA3) — no hardware MX — SGLang converts the weights to block-fp8 `[128,128]` at load and serves them on the tuned ROCm kernels (§2.3). The vision tower stays unquantized.
|
||||||
|
|
||||||
**Recommended generation** (from the model card): `temperature` 1.0, `top_p` 0.95, `top_k` 40 — SGLang applies these automatically from the model's `generation_config.json`.
|
**Recommended generation**: the model's `generation_config.json` sets `temperature` 1.0 / `top_p` 0.95, which SGLang applies automatically (the default `--sampling-defaults model`). The model card additionally suggests `top_k` 40, but that value is **not** in `generation_config.json`, so SGLang does not apply it by default. `top_k` is a per-request sampling parameter (not a launch flag) — set it per call if you want it, e.g. `extra_body={"top_k": 40}` with the OpenAI client.
|
||||||
|
|
||||||
**Resources:** [HuggingFace](https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8) · [MSA kernel](https://github.com/MiniMax-AI/MSA)
|
**Resources:** [HuggingFace](https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8) · [MSA kernel](https://github.com/MiniMax-AI/MSA)
|
||||||
|
|
||||||
@@ -87,7 +87,7 @@ Key characteristics as served by SGLang:
|
|||||||
|
|
||||||
### 2.1 MSA sparse-attention fast path (recommended for Blackwell users)
|
### 2.1 MSA sparse-attention fast path (recommended for Blackwell users)
|
||||||
|
|
||||||
[MiniMax MSA](https://github.com/MiniMax-AI/MSA) (`fmha_sm100`, MIT-licensed) is the recommended Blackwell kernel for M3's main sparse-attention step — faster and more memory-efficient than the built-in Triton fallback. It is purely additive — install it and the recipe above engages it automatically; without it the same recipe still serves on the built-in Triton path. The swap is numerically equivalent (cosine ≥ 0.99999 vs Triton), decode stays CUDA-graph-capturable, prefill TTFT drops ~9–12% at 8K–64K context, and the MSA path survives memory configurations where the Triton path OOMs.
|
[MiniMax MSA](https://github.com/MiniMax-AI/MSA) (`fmha_sm100`, MIT-licensed) is the recommended Blackwell kernel for M3's main sparse-attention step — faster and more memory-efficient than the built-in Triton fallback. **It ships pre-installed in the M3 dev image** (`lmsysorg/sglang:dev-minimax-m3`, also published under the `dev-cu13-minimax-m3` tag), so the Blackwell recipe above engages it automatically with no extra setup — `import fmha_sm100` works out of the box and the kernels JIT-compile on first use. It is otherwise purely additive: on a custom image, install it (below) and the recipe engages it automatically; without it the same recipe still serves on the built-in Triton path. The swap is numerically equivalent (cosine ≥ 0.99999 vs Triton), decode stays CUDA-graph-capturable, prefill TTFT drops ~9–12% at 8K–64K context, and the MSA path survives memory configurations where the Triton path OOMs.
|
||||||
|
|
||||||
**Requirements** (from the [MSA README](https://github.com/MiniMax-AI/MSA#requirements)):
|
**Requirements** (from the [MSA README](https://github.com/MiniMax-AI/MSA#requirements)):
|
||||||
|
|
||||||
@@ -95,10 +95,12 @@ Key characteristics as served by SGLang:
|
|||||||
- **Toolchain**: CUDA Toolkit with `nvcc` ≥ 12.x on `PATH` (or `CUDA_HOME` set) — the kernels are JIT-compiled at first import.
|
- **Toolchain**: CUDA Toolkit with `nvcc` ≥ 12.x on `PATH` (or `CUDA_HOME` set) — the kernels are JIT-compiled at first import.
|
||||||
- **Python**: ≥ 3.10; **OS**: Linux — works on both **x86_64 and aarch64 (Grace, e.g. GB200 / GB300)**; the aarch64 build needs no source edits.
|
- **Python**: ≥ 3.10; **OS**: Linux — works on both **x86_64 and aarch64 (Grace, e.g. GB200 / GB300)**; the aarch64 build needs no source edits.
|
||||||
|
|
||||||
<Accordion title="Install MSA & verify the gate (Python)">
|
<Accordion title="Install MSA (only on a custom image) & verify the gate (Python)">
|
||||||
|
|
||||||
|
The M3 Blackwell dev images above already bundle MSA, so you can skip straight to the gate check. The `git clone` / `pip install` steps are only needed on a custom image that doesn't have `fmha_sm100`.
|
||||||
|
|
||||||
```bash Command
|
```bash Command
|
||||||
# --recursive pulls the CUTLASS submodule required for JIT compilation
|
# Only on a custom image: --recursive pulls the CUTLASS submodule required for JIT compilation
|
||||||
git clone --recursive https://github.com/MiniMax-AI/MSA.git msa
|
git clone --recursive https://github.com/MiniMax-AI/MSA.git msa
|
||||||
cd msa && pip install .
|
cd msa && pip install .
|
||||||
# Verify the SGLang gate (True -> MSA engaged on this device; False -> Triton fallback):
|
# Verify the SGLang gate (True -> MSA engaged on this device; False -> Triton fallback):
|
||||||
@@ -123,11 +125,11 @@ For multimodal (image) serving, keep the same text recipe above — `--attention
|
|||||||
|
|
||||||
### 2.2 Memory and workload tuning
|
### 2.2 Memory and workload tuning
|
||||||
|
|
||||||
The NVIDIA Blackwell recipe is the validated single-node **4-GPU (`--tp 4`)** config, which is also the GB200 / GB300 single-node ceiling. It runs identically on B200 (sm_100), B300 (sm_103), and GB300 (sm_103, aarch64); GB200 (sm_100, aarch64) is inferred-supported — both of its axes are validated above — but not directly benchmarked. The AMD recipes use **8-GPU (`--tp 8`)**.
|
The NVIDIA Blackwell recipes are validated single-node: **B200 at `--tp 8`** and **B300 / GB300 at `--tp 4`** (4-GPU is also the GB200 / GB300 single-node ceiling). GB200 (sm_100, aarch64) is inferred-supported — both of its axes are validated above (B200 is sm_100; GB300 is sm_103 aarch64) — but not directly benchmarked. The AMD recipes use **8-GPU (`--tp 8`)**.
|
||||||
|
|
||||||
- **Memory**: `--mem-fraction-static` trades KV-pool capacity against prefill **activation headroom** — `0.75` is the safe default on NVIDIA (`0.80` on AMD). A higher value is fine at low concurrency but OOMs under high concurrency or long context, so raise it only for interactive single-stream serving.
|
- **Memory**: `--mem-fraction-static` reserves GPU memory for weights + KV pool; the rest is prefill **activation headroom**. The value scales with *free* memory per GPU (card capacity minus per-GPU weight), so it tracks the card more than the TP degree: **`0.65` on B200** (180 GB — less headroom once weights are resident) and **`0.75` on the larger-memory B300 / GB300** (`0.80` on AMD). Lower TP packs more weight per GPU, so a tighter config needs a *lower* value — B200 needs `0.65` even at `--tp 4`. Raising it past the validated value is fine only for low-concurrency single-stream serving; it OOMs under high concurrency or long context.
|
||||||
- **Long context (32K+)**: keep `--mem-fraction-static` at the platform default and raise `--chunked-prefill-size` to `16384`. Decode TPOT stays roughly flat in context length thanks to sparse attention; 1K–128K prompts are validated.
|
- **Long context (32K+)**: keep `--mem-fraction-static` at the platform default and raise `--chunked-prefill-size` to `16384`. Decode TPOT stays roughly flat in context length thanks to sparse attention; 1K–128K prompts are validated.
|
||||||
- **8-GPU nodes**: B200 / B300 hosts with 8 GPUs can use `--tp 8` for more throughput / KV headroom; tp4 is documented as the NVIDIA cross-family common denominator.
|
- **Scaling TP**: B200 is documented at `--tp 8`; B300 / GB200 / GB300 at `--tp 4` (the single-node cross-family common denominator). On an 8-GPU B300 host you can also raise to `--tp 8` for more throughput / KV headroom.
|
||||||
- **Expert parallelism**: to trade latency for throughput add `--ep` (see [Expert Parallelism Deployment](../../../docs/advanced_features/expert_parallelism)). On AMD, set `--ep` equal to `--tp`. Shared-experts fusion is automatically disabled when EP > 1; on AMD standard EP the server also disables `--enable-aiter-allreduce-fusion` automatically to preserve accuracy.
|
- **Expert parallelism**: to trade latency for throughput add `--ep` (see [Expert Parallelism Deployment](../../../docs/advanced_features/expert_parallelism)). On AMD, set `--ep` equal to `--tp`. Shared-experts fusion is automatically disabled when EP > 1; on AMD standard EP the server also disables `--enable-aiter-allreduce-fusion` automatically to preserve accuracy.
|
||||||
- `--trust-remote-code` is required to load the MiniMax config / processor classes.
|
- `--trust-remote-code` is required to load the MiniMax config / processor classes.
|
||||||
|
|
||||||
|
|||||||
@@ -2,26 +2,27 @@
|
|||||||
// minimax-m3.jsx cells. See _deployment.jsx for the speed/accuracy schema.
|
// minimax-m3.jsx cells. See _deployment.jsx for the speed/accuracy schema.
|
||||||
//
|
//
|
||||||
// SPEED — bench_serving --flush-cache, random isl2048/osl256, max_concurrency 64,
|
// SPEED — bench_serving --flush-cache, random isl2048/osl256, max_concurrency 64,
|
||||||
// CUDA graph on. B200 (tp4, MXFP8, MSA fmha_sm100 path) and H200 (tp8, bf16,
|
// CUDA graph on. B200 (tp8, MXFP8, MSA fmha_sm100 path; re-measured 2026-06-15
|
||||||
// built-in Triton sparse) are measured on PR #27944 — warm steady-state from a
|
// with piecewise CUDA graph default-on) and H200 (tp8, bf16, built-in Triton
|
||||||
// 3-run sweep (the cold-start first run, ~2x slower, is excluded). B300 / GB300
|
// sparse) are measured on PR #27944 — warm steady-state from a 3-run sweep (the
|
||||||
// rows are the earlier sglang main (2026-06-11) tp4 MSA
|
// B200 3-run is identical; the H200 cold-start first run, ~2x slower, is
|
||||||
// numbers, pending a #27944 re-measure on their own boxes. GB200 is a bare-match
|
// excluded). B300 / GB300
|
||||||
|
// rows are the earlier 2026-06-11 tp4 MSA numbers (pre-piecewise),
|
||||||
|
// pending a #27944 re-measure on their own boxes. GB200 is a bare-match
|
||||||
// stub (inferred-supported, not benchmarked). AMD: MI355X at 8-GPU tp8 (native
|
// stub (inferred-supported, not benchmarked). AMD: MI355X at 8-GPU tp8 (native
|
||||||
// MXFP8) carries a bench_serving speed row; MI300X (MXFP8 -> block-fp8) was
|
// MXFP8) carries a bench_serving speed row; MI300X (MXFP8 -> block-fp8) was
|
||||||
// accuracy-only. MI350X / MI325X inherit their same-arch sibling's recipe
|
// accuracy-only. MI350X / MI325X inherit their same-arch sibling's recipe
|
||||||
// (stubs). (sgl-eval does NOT measure serving throughput — TTFT/TPOT/tok-s come
|
// (stubs). (sgl-eval does NOT measure serving throughput — TTFT/TPOT/tok-s come
|
||||||
// from sglang.bench_serving.)
|
// from sglang.bench_serving.)
|
||||||
//
|
//
|
||||||
// GSM8K — unified on a SINGLE harness: sgl-eval (github.com/sgl-project/sgl-eval)
|
// GSM8K / GPQA — unified on a SINGLE harness: sgl-eval (github.com/sgl-project/sgl-eval)
|
||||||
// `run gsm8k`, full 1319-question test split, chat endpoint with --thinking
|
// `run gsm8k` (full 1319) / `run gpqa` (GPQA Diamond 198, n-repeats 4), chat
|
||||||
// (M3's reasoning path) + M3's recommended sampling (temp 1.0 / top_p 0.95 /
|
// endpoint with --thinking (M3's reasoning path) + M3's recommended sampling
|
||||||
// top_k 40), symbolic grading. This is the config's Reproduce command. B200
|
// (temp 1.0 / top_p 0.95), symbolic grading. This is the config's Reproduce command.
|
||||||
// (MSA path) and H200 (bf16, built-in Triton sparse) are measured on PR #27944.
|
// H200 is stable at GSM8K 97.04% (std 0.0). B200 was re-measured 2026-06-15 on
|
||||||
// 3-run results: H200 is stable at 97.04% (std 0.0); B200's fresh-server 94.4%
|
// minimax-m3-upstream (piecewise + MSA decode fix): GSM8K 96.51% recommended /
|
||||||
// (= greedy) drifts down over sustained runs interleaved with bench (an
|
// 96.89% greedy (stable single-run), GPQA pass@1[avg-of-4] 89.14% — the merged
|
||||||
// MSA-under-load serving issue under investigation), so it reports the
|
// MSA decode fix resolves the earlier fresh-server-94.4%-then-drift under-load issue.
|
||||||
// fresh-server value, not the drifted mean.
|
|
||||||
// Per-platform re-measurement under sgl-eval is in progress; rows still pending
|
// Per-platform re-measurement under sgl-eval is in progress; rows still pending
|
||||||
// show `gsm8k_pct: null` (no GSM8K row rendered) with the legacy-harness number
|
// show `gsm8k_pct: null` (no GSM8K row rendered) with the legacy-harness number
|
||||||
// kept in a comment. Legacy harnesses were NOT comparable across platforms
|
// kept in a comment. Legacy harnesses were NOT comparable across platforms
|
||||||
@@ -29,14 +30,18 @@
|
|||||||
// which is exactly why we re-measure on one harness.
|
// which is exactly why we re-measure on one harness.
|
||||||
export const benchmarks = [
|
export const benchmarks = [
|
||||||
{
|
{
|
||||||
|
// B200 re-measured 2026-06-15 at tp8 on minimax-m3-upstream (piecewise CUDA
|
||||||
|
// graph default-on + AR-fusion revert/off + MSA decode fix). The earlier
|
||||||
|
// #27944 tp4 speed + GSM8K drift were pre-fix; the merged MSA decode fix
|
||||||
|
// resolves the drift (stable single-run greedy 96.89% / recommended 96.51%).
|
||||||
match: { hw: "b200", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
match: { hw: "b200", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
||||||
sglang_version: "PR #27944",
|
sglang_version: "PR #27944",
|
||||||
speed: [
|
speed: [
|
||||||
// bench_serving --flush-cache, MSA path; warm steady-state (3-run, cold-start run-1 excluded).
|
// bench_serving --flush-cache, MSA path, tp8; warm steady-state (3-run, identical).
|
||||||
{ workload: { dataset: "random", isl: 2048, osl: 256, max_concurrency: 64, num_prompts: 128 },
|
{ workload: { dataset: "random", isl: 2048, osl: 256, max_concurrency: 64, num_prompts: 128 },
|
||||||
ttft_ms: 749, tpot_ms: 61.5, tokens_per_sec_per_gpu: 249 },
|
ttft_ms: 1580, tpot_ms: 24.1, tokens_per_sec_per_gpu: 265 },
|
||||||
],
|
],
|
||||||
accuracy: { gsm8k_pct: 94.4 }, // #27944, sgl-eval --thinking, full 1319, recommended sampling (temp 1.0/top_p 0.95/top_k 40), MSA path; fresh-server 94.4% (greedy 94.16%; --no-thinking 88.6%). NOTE: 3 sustained runs interleaved with bench drifted 94.4->89.2->86.2 — an MSA-under-load serving issue (under investigation), not the model accuracy.
|
accuracy: { gpqa_pct: 89.1, gsm8k_pct: 96.5 }, // 2026-06-15, sgl-eval --thinking, recommended sampling (temp 1.0/top_p 0.95), tp8. GSM8K full 1319 = 96.51% (greedy 96.89%). GPQA Diamond 198, n-repeats 4 = pass@1[avg-of-4] 89.14% +/-1.73% (pass@4 95.45%, majority@4 93.52%).
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
// Hopper H200: bf16 build (MXFP8 is Blackwell-only) at tp8, built-in Triton
|
// Hopper H200: bf16 build (MXFP8 is Blackwell-only) at tp8, built-in Triton
|
||||||
@@ -52,7 +57,7 @@ export const benchmarks = [
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
match: { hw: "b300", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
match: { hw: "b300", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
||||||
sglang_version: "main (2026-06-11)",
|
sglang_version: "PR #27944",
|
||||||
speed: [
|
speed: [
|
||||||
{ workload: { dataset: "random", isl: 2048, osl: 256, max_concurrency: 64 },
|
{ workload: { dataset: "random", isl: 2048, osl: 256, max_concurrency: 64 },
|
||||||
ttft_ms: null, tpot_ms: 32.8, tokens_per_sec_per_gpu: 365 },
|
ttft_ms: null, tpot_ms: 32.8, tokens_per_sec_per_gpu: 365 },
|
||||||
@@ -63,7 +68,7 @@ export const benchmarks = [
|
|||||||
{ match: { hw: "gb200", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" } },
|
{ match: { hw: "gb200", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" } },
|
||||||
{
|
{
|
||||||
match: { hw: "gb300", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
match: { hw: "gb300", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
||||||
sglang_version: "main (2026-06-11)",
|
sglang_version: "PR #27944",
|
||||||
speed: [
|
speed: [
|
||||||
{ workload: { dataset: "random", isl: 2048, osl: 256, max_concurrency: 64 },
|
{ workload: { dataset: "random", isl: 2048, osl: 256, max_concurrency: 64 },
|
||||||
ttft_ms: 4746, tpot_ms: 39.3, tokens_per_sec_per_gpu: 277 },
|
ttft_ms: 4746, tpot_ms: 39.3, tokens_per_sec_per_gpu: 277 },
|
||||||
@@ -77,7 +82,7 @@ export const benchmarks = [
|
|||||||
// No TTFT/TPOT reported for this run.
|
// No TTFT/TPOT reported for this run.
|
||||||
{
|
{
|
||||||
match: { hw: "mi355x", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
match: { hw: "mi355x", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
||||||
sglang_version: "main (2026-06-11)",
|
sglang_version: "PR #27944",
|
||||||
speed: [
|
speed: [
|
||||||
{ workload: { dataset: "random", isl: 1024, osl: 1024, max_concurrency: 64, num_prompts: 640 },
|
{ workload: { dataset: "random", isl: 1024, osl: 1024, max_concurrency: 64, num_prompts: 640 },
|
||||||
ttft_ms: null, tpot_ms: null, tokens_per_sec_per_gpu: 210 },
|
ttft_ms: null, tpot_ms: null, tokens_per_sec_per_gpu: 210 },
|
||||||
@@ -89,7 +94,7 @@ export const benchmarks = [
|
|||||||
// MI300X (gfx942): MXFP8 -> block-fp8 [128,128].
|
// MI300X (gfx942): MXFP8 -> block-fp8 [128,128].
|
||||||
{
|
{
|
||||||
match: { hw: "mi300x", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
match: { hw: "mi300x", variant: "default", quant: "mxfp8", strategy: "balanced", nodes: "single" },
|
||||||
sglang_version: "main (2026-06-11)",
|
sglang_version: "PR #27944",
|
||||||
accuracy: { gsm8k_pct: null }, // TODO: pending sgl-eval re-measure on MI300X (legacy run_eval 1319: 92.0, triton 0.917-0.929 / aiter ~0.929)
|
accuracy: { gsm8k_pct: null }, // TODO: pending sgl-eval re-measure on MI300X (legacy run_eval 1319: 92.0, triton 0.917-0.929 / aiter ~0.929)
|
||||||
},
|
},
|
||||||
// MI325X (gfx942): inferred-supported from MI300X, not separately benchmarked.
|
// MI325X (gfx942): inferred-supported from MI300X, not separately benchmarked.
|
||||||
|
|||||||
@@ -64,11 +64,19 @@ sgl-eval run gsm8k \\
|
|||||||
--model {{MODEL_NAME}} \\
|
--model {{MODEL_NAME}} \\
|
||||||
--temperature 1.0 --top-p 0.95 \\
|
--temperature 1.0 --top-p 0.95 \\
|
||||||
--thinking`,
|
--thinking`,
|
||||||
|
gpqa_pct:
|
||||||
|
`pip install git+https://github.com/sgl-project/sgl-eval
|
||||||
|
sgl-eval run gpqa \\
|
||||||
|
--base-url http://{{CURL_HOST}}:{{CURL_PORT}}/v1 \\
|
||||||
|
--model {{MODEL_NAME}} \\
|
||||||
|
--temperature 1.0 --top-p 0.95 \\
|
||||||
|
--thinking --n-repeats 4 --max-tokens 40960`,
|
||||||
},
|
},
|
||||||
numPromptsByConc: { 24: 24, 64: 128 },
|
numPromptsByConc: { 24: 24, 64: 128 },
|
||||||
},
|
},
|
||||||
|
|
||||||
accuracyLabels: [
|
accuracyLabels: [
|
||||||
|
["gpqa_pct", "GPQA Diamond", "%"],
|
||||||
["gsm8k_pct", "GSM8K", "%"],
|
["gsm8k_pct", "GSM8K", "%"],
|
||||||
],
|
],
|
||||||
|
|
||||||
@@ -173,10 +181,12 @@ sgl-eval run gsm8k \\
|
|||||||
},
|
},
|
||||||
},
|
},
|
||||||
|
|
||||||
// NVIDIA Blackwell: one validated single-node tp4 recipe per family. fa4 +
|
// NVIDIA Blackwell: one validated single-node recipe per family — tp4 across
|
||||||
// page 128 + deep_gemm are the M3 SM100 auto-defaults on current main, so this
|
// B300 / GB200 / GB300, tp8 on B200. fa4 + page 128 + deep_gemm are the M3
|
||||||
// is also the bare-launch behavior; they engage MiniMax's MSA sparse-attention
|
// SM100 auto-defaults on current main, so this is also the bare-launch
|
||||||
// kernel when fmha_sm100 is installed (see Configuration Tips), Triton otherwise.
|
// behavior; they engage MiniMax's MSA sparse-attention kernel (fmha_sm100,
|
||||||
|
// pre-installed in the dev-minimax-m3 images; see Configuration Tips), Triton
|
||||||
|
// fallback otherwise.
|
||||||
// AMD: tp8. MI350X/MI355X (gfx950) serve MXFP8 natively (backends auto). MI300X/
|
// AMD: tp8. MI350X/MI355X (gfx950) serve MXFP8 natively (backends auto). MI300X/
|
||||||
// MI325X (gfx942) need --attention-backend aiter + --moe-runner-backend triton,
|
// MI325X (gfx942) need --attention-backend aiter + --moe-runner-backend triton,
|
||||||
// and the MXFP8 weights are auto-converted to block-fp8 at load; the cold-start
|
// and the MXFP8 weights are auto-converted to block-fp8 at load; the cold-start
|
||||||
@@ -191,12 +201,12 @@ sgl-eval run gsm8k \\
|
|||||||
"--model-path {{MODEL_NAME}}",
|
"--model-path {{MODEL_NAME}}",
|
||||||
"--reasoning-parser auto",
|
"--reasoning-parser auto",
|
||||||
"--tool-call-parser auto",
|
"--tool-call-parser auto",
|
||||||
"--tp 4",
|
"--tp 8",
|
||||||
"--attention-backend fa4",
|
"--attention-backend fa4",
|
||||||
"--page-size 128",
|
"--page-size 128",
|
||||||
"--moe-runner-backend deep_gemm",
|
"--moe-runner-backend deep_gemm",
|
||||||
"--chunked-prefill-size 8192",
|
"--chunked-prefill-size 8192",
|
||||||
"--mem-fraction-static 0.75",
|
"--mem-fraction-static 0.65",
|
||||||
"--host {{HOST_IP}}",
|
"--host {{HOST_IP}}",
|
||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
],
|
],
|
||||||
|
|||||||
Reference in New Issue
Block a user