docs(cookbook): re-benchmark DeepSeek-V4 on sglang 0.5.15 (#31363)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
This commit is contained in:
Douglas Yang
2026-07-21 22:55:36 +00:00
committed by GitHub
co-authored by Claude Opus 4.8 Zijie Xia
parent 9057db9417
commit 4a55fdba0b
8 changed files with 231 additions and 71 deletions
@@ -177,7 +177,8 @@ One entry per measured block only (cells without entries already render
total (in+out) tok/s/GPU = `output tok/s ÷ (tp × nnodes) ×
(isl+osl)/osl` — stored directly (the card shows it as-is). TTFT/TPOT
take the P50 (median) rows; set `config.latencyPercentile` (default `"P50"`; use
`"Mean"` only for legacy Mean-recorded data — temporary, being migrated to P50).
`"Mean"` only for legacy Mean-recorded data — temporary, being migrated to P50; an
entry-level `latencyPercentile` overrides the page value per cell).
Put the workload's
`num_prompts` into `workload`. **`config.accuracyLabels` is required whenever
the benchmarks carry accuracy data** — the engine ships no default eval set