Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
62a4a6ea0e
commit
e4008de757
@@ -2,7 +2,6 @@
|
|||||||
title: MiniCPM-V 4.6
|
title: MiniCPM-V 4.6
|
||||||
metatags:
|
metatags:
|
||||||
description: "Deploy OpenBMB MiniCPM-V 4.6 (Qwen3.5-style hybrid GDN backbone + NaViT vision encoder) on NVIDIA GPUs with SGLang — multimodal text + image + video, slicing for high-resolution images."
|
description: "Deploy OpenBMB MiniCPM-V 4.6 (Qwen3.5-style hybrid GDN backbone + NaViT vision encoder) on NVIDIA GPUs with SGLang — multimodal text + image + video, slicing for high-resolution images."
|
||||||
tag: NEW
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 1. Model Introduction
|
## 1. Model Introduction
|
||||||
|
|||||||
@@ -0,0 +1,220 @@
|
|||||||
|
---
|
||||||
|
title: MiniCPM5-2B
|
||||||
|
description: "Deploy OpenBMB MiniCPM5-2B with SGLang — a 2.5B dense Llama-architecture on-device model with 131K context, thinking mode, XML tool calling and DSpark speculative decoding on H200, RTX PRO 6000, RTX 5090 and DGX Spark."
|
||||||
|
tag: NEW
|
||||||
|
---
|
||||||
|
|
||||||
|
## Deployment
|
||||||
|
|
||||||
|
<a id="install" />
|
||||||
|
|
||||||
|
<Accordion title="Install SGLang">
|
||||||
|
|
||||||
|
For all methods and hardware platforms, see the [official SGLang installation guide](../../../docs/get-started/install). The two paths below match the **Python / Docker** toggle in the command panel.
|
||||||
|
|
||||||
|
<Tabs>
|
||||||
|
|
||||||
|
<Tab title="Python (pip / uv)">
|
||||||
|
|
||||||
|
```bash Command
|
||||||
|
pip install --upgrade pip
|
||||||
|
pip install uv
|
||||||
|
uv pip install --prerelease=allow "sglang>=0.5.12"
|
||||||
|
```
|
||||||
|
|
||||||
|
Then run the **Python** output of the command panel below in that environment.
|
||||||
|
|
||||||
|
</Tab>
|
||||||
|
|
||||||
|
<Tab title="Docker">
|
||||||
|
|
||||||
|
```bash Command
|
||||||
|
docker pull lmsysorg/sglang:dev
|
||||||
|
```
|
||||||
|
|
||||||
|
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
|
||||||
|
|
||||||
|
</Tab>
|
||||||
|
|
||||||
|
</Tabs>
|
||||||
|
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
Pick your card to generate the launch command. MiniCPM5-2B is a 2.5B dense model and runs single-GPU at TP=1 on every supported card, so the page ships one operating point per card plus an optional **DSpark** speculative-decoding row.
|
||||||
|
|
||||||
|
import { Deployment } from "/src/snippets/_deployment.jsx";
|
||||||
|
import { config } from "/src/snippets/configs/openbmb/minicpm5-2b.jsx";
|
||||||
|
import { benchmarks } from "/src/snippets/configs/openbmb/minicpm5-2b-benchmarks.jsx";
|
||||||
|
|
||||||
|
<Deployment config={config} benchmarks={benchmarks} />
|
||||||
|
|
||||||
|
<Note>
|
||||||
|
Speed numbers exist for the RTX 5090 and DGX Spark cells — a single card at
|
||||||
|
random 1024/1024, recorded as Mean. The H200 and RTX PRO 6000 cells are
|
||||||
|
pending measurement, and no accuracy numbers have been taken on any platform
|
||||||
|
yet. The DSpark overlay carries no speed numbers on any card.
|
||||||
|
</Note>
|
||||||
|
|
||||||
|
## Playground
|
||||||
|
|
||||||
|
The Playground is where you experiment with **SGLang features beyond the recipes above**. The Deploy panel emits this model's documented launch recipes; the Playground lets you turn on additional knobs on top of whichever cell the Deploy panel is currently showing.
|
||||||
|
|
||||||
|
import { Playground } from "/src/snippets/_playground.jsx";
|
||||||
|
|
||||||
|
<Playground config={config} />
|
||||||
|
|
||||||
|
## 1. Model Introduction
|
||||||
|
|
||||||
|
**MiniCPM5-2B** is the second model in OpenBMB's MiniCPM5 series, following MiniCPM5-1B. It is a dense 2B-class Transformer built for on-device assistants, local deployment, coding agents and tool-use workflows — scenarios where a compact model with a small deployment footprint is preferred. OpenBMB reports 2B-class open-source SOTA within its comparison set (average 53.9), with its clearest advantages in code reasoning, math reasoning, long-context understanding, tool use and agentic tasks.
|
||||||
|
|
||||||
|
The checkpoint uses the standard `LlamaForCausalLM` architecture — no custom kernels and no model-code fork — so SGLang loads it through the stock Llama path. It is 2,516,756,480 parameters (1,981,982,720 non-embedding) over 42 layers with GQA (16 query heads, 2 KV heads), and a native context length of 131,072 tokens. Post-training runs SFT → RL → On-Policy Distillation, which merges 16 RL expert models into the single released checkpoint. Weights are released under the [Apache-2.0](https://github.com/OpenBMB/MiniCPM/blob/main/LICENSE) license.
|
||||||
|
|
||||||
|
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||||
|
<colgroup>
|
||||||
|
<col style={{width: "34%"}} />
|
||||||
|
<col style={{width: "22%"}} />
|
||||||
|
<col style={{width: "44%"}} />
|
||||||
|
</colgroup>
|
||||||
|
<thead>
|
||||||
|
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||||
|
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Checkpoint</th>
|
||||||
|
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Precision</th>
|
||||||
|
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Role on this page</th>
|
||||||
|
</tr>
|
||||||
|
</thead>
|
||||||
|
<tbody>
|
||||||
|
<tr>
|
||||||
|
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><a href="https://huggingface.co/openbmb/MiniCPM5-2B">openbmb/MiniCPM5-2B</a></td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>BF16</td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>The served model in every cell above (final release, post-trained with RL + OPD).</td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.05)"}}><a href="https://huggingface.co/openbmb/MiniCPM5-2B-DSpark">openbmb/MiniCPM5-2B-DSpark</a></td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>BF16 draft</td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Draft model loaded by the <strong>DSPARK</strong> row of the Deploy panel; not served on its own.</td>
|
||||||
|
</tr>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
|
||||||
|
OpenBMB also publishes SFT-only, mid-training and base checkpoints, plus GGUF, MLX and GPTQ-Int4 exports for llama.cpp / Ollama / LM Studio / Apple Silicon. Those target other runtimes and are not part of the SGLang matrix above.
|
||||||
|
|
||||||
|
**Recommended generation:** `temperature=1.0`, `top_p=0.95` (informational — SGLang reads the checkpoint's `generation_config.json`; do not hardcode these in client code).
|
||||||
|
|
||||||
|
**Resources:** [HuggingFace](https://huggingface.co/openbmb/MiniCPM5-2B) · [ModelScope](https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B) · [GitHub](https://github.com/OpenBMB/MiniCPM) · [Tech report](https://arxiv.org/pdf/2506.07900).
|
||||||
|
|
||||||
|
## 2. Advanced Usage
|
||||||
|
|
||||||
|
### 2.1 Thinking Mode
|
||||||
|
|
||||||
|
Thinking is controlled by the chat template's `enable_thinking` flag, passed per request through `chat_template_kwargs`. The `qwen3` reasoning parser — already in every generated command, and toggleable from the **Parsers** card in the [Playground above](#playground) — splits the `<think>` segment into `message.reasoning_content` and leaves the final answer in `message.content`. Drop the flag and the thinking text stays inline in `content`, closing tag and all.
|
||||||
|
|
||||||
|
<Accordion title="Thinking On / Off (Python)">
|
||||||
|
|
||||||
|
```python Example
|
||||||
|
from openai import OpenAI
|
||||||
|
|
||||||
|
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
|
||||||
|
|
||||||
|
# Thinking on (chat-template default).
|
||||||
|
resp = client.chat.completions.create(
|
||||||
|
model="openbmb/MiniCPM5-2B",
|
||||||
|
messages=[{"role": "user", "content": "What is 15% of 240?"}],
|
||||||
|
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
|
||||||
|
)
|
||||||
|
msg = resp.choices[0].message
|
||||||
|
print("Reasoning:", getattr(msg, "reasoning_content", None))
|
||||||
|
print("Answer:", msg.content)
|
||||||
|
|
||||||
|
# Thinking off — a direct answer, no deliberation segment.
|
||||||
|
resp = client.chat.completions.create(
|
||||||
|
model="openbmb/MiniCPM5-2B",
|
||||||
|
messages=[{"role": "user", "content": "What is 15% of 240?"}],
|
||||||
|
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
|
||||||
|
)
|
||||||
|
print("Without thinking:", resp.choices[0].message.content)
|
||||||
|
```
|
||||||
|
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="Example Output">
|
||||||
|
|
||||||
|
```text Output
|
||||||
|
Reasoning: We are asked: "What is 15% of 240?" This is a simple percentage calculation. To find 15% of 240, we multiply 240 by 15% (which is 0.15). So: 240 × 0.15 = 36. Alternatively, we can think of it as (15/100) × 240 = (15 × 240)/100 = 3600/100 = 36. So the answer is 36.
|
||||||
|
|
||||||
|
We need to respond in a helpful way. The user might be testing or seeking quick answer. Let's provide a clear response with explanation if needed, but since it's straightforward, we can just give the answer directly or briefly explain.
|
||||||
|
|
||||||
|
Answer:
|
||||||
|
|
||||||
|
15% of 240 is **36**.
|
||||||
|
|
||||||
|
To calculate:
|
||||||
|
\( 240 \times 0.15 = 36 \)
|
||||||
|
Reasoning: None
|
||||||
|
Answer: To find 15% of 240, multiply 240 by 15% (which is 0.15):
|
||||||
|
|
||||||
|
\[
|
||||||
|
240 \times 0.15 = 36
|
||||||
|
\]
|
||||||
|
|
||||||
|
So, 15% of 240 is **36**.
|
||||||
|
```
|
||||||
|
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
### 2.2 Tool Calling
|
||||||
|
|
||||||
|
MiniCPM5-2B emits XML-style tool calls (`<function name="..."><param name="...">...</param></function>`), and SGLang's built-in `minicpm5` detector converts them to OpenAI-compatible `tool_calls`, parallel calls included. The parser is already in every command the Deploy panel generates; the **Tool Call Parser** chip in the Playground's **Parsers** card is an opt-out, not an opt-in.
|
||||||
|
|
||||||
|
<Accordion title="Tool Calling (Python)">
|
||||||
|
|
||||||
|
```python Example
|
||||||
|
from openai import OpenAI
|
||||||
|
|
||||||
|
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
|
||||||
|
|
||||||
|
tools = [{
|
||||||
|
"type": "function",
|
||||||
|
"function": {
|
||||||
|
"name": "get_weather",
|
||||||
|
"description": "Get the current weather for a city.",
|
||||||
|
"parameters": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"city": {"type": "string", "description": "City name"},
|
||||||
|
"date": {"type": "string", "description": "YYYY-MM-DD"},
|
||||||
|
},
|
||||||
|
"required": ["city"],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
}]
|
||||||
|
|
||||||
|
resp = client.chat.completions.create(
|
||||||
|
model="openbmb/MiniCPM5-2B",
|
||||||
|
messages=[{"role": "user", "content": "What is the weather in Beijing today?"}],
|
||||||
|
tools=tools,
|
||||||
|
)
|
||||||
|
|
||||||
|
msg = resp.choices[0].message
|
||||||
|
# With the qwen3 reasoning parser on, the thinking segment lands in
|
||||||
|
# `reasoning_content` and `content` may be empty on a tool-call turn.
|
||||||
|
print("Reasoning:", getattr(msg, "reasoning_content", None))
|
||||||
|
print("Content:", msg.content)
|
||||||
|
for call in (msg.tool_calls or []):
|
||||||
|
print("Tool:", call.function.name, call.function.arguments)
|
||||||
|
```
|
||||||
|
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="Example Output">
|
||||||
|
|
||||||
|
```text Output
|
||||||
|
Reasoning:
|
||||||
|
The user is asking about the weather in Beijing today. I have access to a get_weather function that can help with this. Let me call it to get the current weather for Beijing.
|
||||||
|
|
||||||
|
Content:
|
||||||
|
|
||||||
|
|
||||||
|
Tool: get_weather {"city": "Beijing"}
|
||||||
|
```
|
||||||
|
|
||||||
|
</Accordion>
|
||||||
@@ -136,7 +136,7 @@ metatags:
|
|||||||
<Card
|
<Card
|
||||||
title="OpenBMB"
|
title="OpenBMB"
|
||||||
mode="card"
|
mode="card"
|
||||||
href="/cookbook/autoregressive/OpenBMB/MiniCPM-V-4_6"
|
href="/cookbook/autoregressive/OpenBMB/MiniCPM5-2B"
|
||||||
img="/cards/logos/openbmb.png"
|
img="/cards/logos/openbmb.png"
|
||||||
/>
|
/>
|
||||||
<Card
|
<Card
|
||||||
|
|||||||
@@ -1404,6 +1404,7 @@
|
|||||||
{
|
{
|
||||||
"group": "OpenBMB",
|
"group": "OpenBMB",
|
||||||
"pages": [
|
"pages": [
|
||||||
|
"cookbook/autoregressive/OpenBMB/MiniCPM5-2B",
|
||||||
"cookbook/autoregressive/OpenBMB/MiniCPM-V-4_6"
|
"cookbook/autoregressive/OpenBMB/MiniCPM-V-4_6"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
|||||||
@@ -0,0 +1,55 @@
|
|||||||
|
// One entry per cell `match` tuple (same keys as the config cells).
|
||||||
|
//
|
||||||
|
// Only the RTX 5090 row carries numbers. They were measured on a single
|
||||||
|
// RTX 5090 32GB (`zijiexia-5090-1gpu`, `lmsysorg/sglang:dev` @ 30705c004c) with
|
||||||
|
// `python3 -m sglang.bench_serving`, random 1024/1024, against the exact cell
|
||||||
|
// command on this page — including `--mem-fraction-static 0.75
|
||||||
|
// --cuda-graph-max-bs 128`, which is what keeps concurrency 64 and 128
|
||||||
|
// graph-backed (see the model page's Configuration Tips).
|
||||||
|
//
|
||||||
|
// `tokens_per_sec_per_gpu` is total (input+output) per GPU = the measured
|
||||||
|
// output throughput x 2 at 1024/1024 on one GPU.
|
||||||
|
//
|
||||||
|
// Two operating points only: concurrency 1 (single-user latency) and 128 (the
|
||||||
|
// saturated end of the sweep). The intermediate 8 / 32 / 64 rows were measured
|
||||||
|
// but are not published here.
|
||||||
|
//
|
||||||
|
// The source runs did not restate a percentile, so the values are recorded as
|
||||||
|
// Mean (bench_serving's headline lines) rather than claimed as P50 — same
|
||||||
|
// convention as the LFM2.5 page. TTFT was not captured at concurrency 128, so
|
||||||
|
// that cell carries `null` rather than a back-filled guess.
|
||||||
|
//
|
||||||
|
// No accuracy rows: the only GSM8K numbers taken so far came from a custom
|
||||||
|
// chat-template harness, not the `sgl-eval run gsm8k` command the Reproduce
|
||||||
|
// modal would show, and the two are not comparable.
|
||||||
|
|
||||||
|
export const benchmarks = [
|
||||||
|
{
|
||||||
|
match: { hw: "rtx5090", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
|
sglang_version: "dev @ 30705c004c",
|
||||||
|
latencyPercentile: "Mean",
|
||||||
|
speed: [
|
||||||
|
{ workload: { dataset: "random", isl: 1024, osl: 1024, max_concurrency: 1 },
|
||||||
|
ttft_ms: 34, tpot_ms: 4.0, tokens_per_sec_per_gpu: 496 },
|
||||||
|
{ workload: { dataset: "random", isl: 1024, osl: 1024, max_concurrency: 128 },
|
||||||
|
ttft_ms: null, tpot_ms: 11.8, tokens_per_sec_per_gpu: 19280 },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
// Single DGX Spark (GB10), same image and commit as the RTX 5090 row.
|
||||||
|
// TTFT is Mean; at concurrency 64 the Median is 246 ms, because the first
|
||||||
|
// wave of requests all queue behind one prefill.
|
||||||
|
match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
|
sglang_version: "dev @ 30705c004c",
|
||||||
|
latencyPercentile: "Mean",
|
||||||
|
speed: [
|
||||||
|
{ workload: { dataset: "random", isl: 1024, osl: 1024, max_concurrency: 1 },
|
||||||
|
ttft_ms: 85, tpot_ms: 27.7, tokens_per_sec_per_gpu: 72 },
|
||||||
|
{ workload: { dataset: "random", isl: 1024, osl: 1024, max_concurrency: 64 },
|
||||||
|
ttft_ms: 1373, tpot_ms: 42.9, tokens_per_sec_per_gpu: 2892 },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
// Pending — no numbers taken on these platforms yet.
|
||||||
|
{ match: { hw: "h200", variant: "default", quant: "bf16", nodes: "single" } },
|
||||||
|
{ match: { hw: "rtx6000", variant: "default", quant: "bf16", nodes: "single" } },
|
||||||
|
];
|
||||||
@@ -0,0 +1,234 @@
|
|||||||
|
// Single `export const config` literal — no spreads/calls/IIFE (Mintlify re-evals at hydration).
|
||||||
|
// Cells are denormalized: no `--nnodes`/`--node-rank`/`--dist-init-addr`/`--host`/`--port` literals — engine injects them.
|
||||||
|
//
|
||||||
|
// MiniCPM5-2B: 2.5B dense `LlamaForCausalLM` (42 layers, GQA 16Q/2KV, 131072
|
||||||
|
// context). Standard architecture — no custom kernels and no model-code fork —
|
||||||
|
// so every supported card runs it single-GPU at TP=1 with the stock backend.
|
||||||
|
// That leaves one recipe per card, hence a single variant / quantization /
|
||||||
|
// strategy / node option and no parallelism flags in any cell.
|
||||||
|
//
|
||||||
|
// Recipes are the OpenBMB model card's SGLang commands, rewritten from
|
||||||
|
// `python -m sglang.launch_server` to `sglang serve`, plus the parser pair the
|
||||||
|
// model needs to be usable through the OpenAI API. Both are baked into every
|
||||||
|
// cell, so the Parsers card in the Playground reads as an opt-OUT:
|
||||||
|
// --tool-call-parser minicpm5 the model emits XML-style
|
||||||
|
// `<function name="f"><param name="p">v</param></function>`; without the
|
||||||
|
// detector `tool_calls` comes back None and the XML lands in `content`.
|
||||||
|
// --reasoning-parser qwen3 the chat template is Qwen-style
|
||||||
|
// (`<|im_start|>` + `<think>`) and there is no `minicpm5` reasoning
|
||||||
|
// detector, so `qwen3` is the one that applies; without it `</think>`
|
||||||
|
// leaks into `content`.
|
||||||
|
//
|
||||||
|
// DSpark is the separately published draft checkpoint
|
||||||
|
// (openbmb/MiniCPM5-2B-DSpark). It is orthogonal to the card grid, so it is an
|
||||||
|
// overlay row rather than a match dim. No DSpark speed numbers are published:
|
||||||
|
// the speedup tracks acceptance length, which moves with the prompt
|
||||||
|
// distribution, and a random-token dataset inflates it above real traffic.
|
||||||
|
|
||||||
|
export const config = {
|
||||||
|
modelName: "MiniCPM5-2B",
|
||||||
|
|
||||||
|
supportedHardware: ["h200", "rtx6000", "rtx5090", "dgx-spark"],
|
||||||
|
|
||||||
|
// RTX PRO 6000 and RTX 5090 (SM120 / Blackwell workstation + desktop) are not
|
||||||
|
// datacenter parts, so the shared HARDWARE_CATALOG in _deployment.jsx does not
|
||||||
|
// carry them. Ids/labels match the DeepSeek-V4 and Qwen3.8-27B configs.
|
||||||
|
hardware: [
|
||||||
|
{ id: "rtx6000", label: "RTX PRO 6000", vram: "96GB", vendor: "blackwell" },
|
||||||
|
{ id: "rtx5090", label: "RTX 5090", vram: "32GB", vendor: "blackwell" },
|
||||||
|
],
|
||||||
|
|
||||||
|
matchDims: [
|
||||||
|
{ id: "variant", title: "Model Variant", options: [
|
||||||
|
{ id: "default", label: "Default" },
|
||||||
|
] },
|
||||||
|
{ id: "quant", title: "Quantization", options: [
|
||||||
|
{ id: "bf16", label: "BF16" },
|
||||||
|
] },
|
||||||
|
{ id: "nodes", title: "Nodes", options: [
|
||||||
|
{ id: "single", label: "Single Node" },
|
||||||
|
] },
|
||||||
|
],
|
||||||
|
|
||||||
|
overlayDims: [
|
||||||
|
{
|
||||||
|
id: "spec",
|
||||||
|
title: "Speculative Decoding",
|
||||||
|
default: "none",
|
||||||
|
options: [
|
||||||
|
{ id: "none", label: "None" },
|
||||||
|
{
|
||||||
|
id: "dspark", label: "DSPARK",
|
||||||
|
// Verbatim from the model card's DSpark command, including
|
||||||
|
// `--trust-remote-code`: the base checkpoint is plain Llama and does
|
||||||
|
// not need it, the draft checkpoint's config does. gamma = 7, so the
|
||||||
|
// verify window is 8 tokens.
|
||||||
|
flags: [
|
||||||
|
"--trust-remote-code",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
|
"--speculative-draft-model-path openbmb/MiniCPM5-2B-DSpark",
|
||||||
|
"--speculative-dspark-block-size 7",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
|
||||||
|
modelNames: {
|
||||||
|
"default|bf16": "openbmb/MiniCPM5-2B",
|
||||||
|
},
|
||||||
|
|
||||||
|
placeholders: {
|
||||||
|
HOST_IP: { target: "command", label: "Bind host", default: "0.0.0.0" },
|
||||||
|
PORT: { target: "command", label: "Bind port", default: "30000" },
|
||||||
|
HF_TOKEN: { target: "command", label: "HF token (Docker)", default: "<your-hf-token>" },
|
||||||
|
CURL_HOST: { target: "curl", label: "Server host", default: "localhost" },
|
||||||
|
CURL_PORT: { target: "curl", label: "Server port", default: "30000" },
|
||||||
|
},
|
||||||
|
|
||||||
|
curl: `curl http://{{CURL_HOST}}:{{CURL_PORT}}/v1/chat/completions \\
|
||||||
|
-H 'Content-Type: application/json' \\
|
||||||
|
-d '{ "model": "{{MODEL_NAME}}", "messages": [{"role":"user","content":"Who are you? Please briefly introduce yourself."}] }'`,
|
||||||
|
|
||||||
|
// Reproduce command for the Benchmark card's "⚡ Reproduce" modal. No
|
||||||
|
// `accuracy` entry: the page carries no accuracy numbers yet.
|
||||||
|
benchmarkCommands: {
|
||||||
|
speed:
|
||||||
|
`python3 -m sglang.bench_serving \\
|
||||||
|
--backend sglang \\
|
||||||
|
--host {{CURL_HOST}} --port {{CURL_PORT}} \\
|
||||||
|
--model {{MODEL_NAME}} \\
|
||||||
|
--dataset-name {{DATASET}} \\
|
||||||
|
--random-input-len {{ISL}} --random-output-len {{OSL}} \\
|
||||||
|
--num-prompts {{NUM_PROMPTS}} --max-concurrency {{MAX_CONCURRENCY}} \\
|
||||||
|
--flush-cache`,
|
||||||
|
numPromptsByConc: { 1: 10, 128: 512 },
|
||||||
|
},
|
||||||
|
|
||||||
|
// MiniCPM5 support (the `minicpm5` tool-call parser and the DSPARK draft
|
||||||
|
// worker) ships in the SGLang dev image.
|
||||||
|
dockerImages: {
|
||||||
|
h200: "lmsysorg/sglang:dev",
|
||||||
|
rtx6000: "lmsysorg/sglang:dev",
|
||||||
|
rtx5090: "lmsysorg/sglang:dev",
|
||||||
|
"dgx-spark": "lmsysorg/sglang:dev",
|
||||||
|
},
|
||||||
|
|
||||||
|
// Pre-selects the issue template's `model` field on "Submit verified cell".
|
||||||
|
github: {
|
||||||
|
cookbookModel: "openbmb/MiniCPM5-2B",
|
||||||
|
},
|
||||||
|
|
||||||
|
playgroundFeatures: {
|
||||||
|
// The model fits one GPU on every supported card, so TP=1 is the verified
|
||||||
|
// shape; TP=2 is exposed for experimentation only.
|
||||||
|
attention: {
|
||||||
|
knobs: [
|
||||||
|
{ id: "tp", label: "TP", values: [null, 1, 2] },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
|
||||||
|
// ----- Card: "Parsers" -----
|
||||||
|
// Opt-OUT: both flags are already in every cell, so the handler derives
|
||||||
|
// each chip as on and strips the flag when one is toggled off. The
|
||||||
|
// reasoning slug is `qwen3`, not `minicpm5` — see the header note.
|
||||||
|
parsers: {
|
||||||
|
items: [
|
||||||
|
{ id: "reasoning", label: "Reasoning Parser", flag: "--reasoning-parser qwen3" },
|
||||||
|
{ id: "toolCall", label: "Tool Call Parser", flag: "--tool-call-parser minicpm5" },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
|
||||||
|
// ----- Card: "Speculative Decoding" -----
|
||||||
|
// Same DSpark flags as the Deploy panel's overlay row, so the two paths
|
||||||
|
// compose an identical command.
|
||||||
|
speculative: {
|
||||||
|
options: [
|
||||||
|
{ id: "current", label: "Inherited from base" },
|
||||||
|
{ id: "off", label: "Off (greedy)" },
|
||||||
|
{ id: "dspark", label: "DSpark",
|
||||||
|
flags: ["--trust-remote-code",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
|
"--speculative-draft-model-path openbmb/MiniCPM5-2B-DSpark",
|
||||||
|
"--speculative-dspark-block-size 7"] },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
|
||||||
|
// One recipe per card — the model card's SGLang launch line, plus the
|
||||||
|
// `minicpm5` tool-call parser it recommends for agent workloads and the
|
||||||
|
// `qwen3` reasoning parser its `<think>` template needs.
|
||||||
|
cells: [
|
||||||
|
{
|
||||||
|
match: { hw: "h200", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
|
verified: true,
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser minicpm5",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
// Verification round still open on this card. `verificationStatus` alone,
|
||||||
|
// with NO `verified: true` baseline: the boolean is what the Playground
|
||||||
|
// reads for its own badge, so leaving it on would make the Playground
|
||||||
|
// claim "Verified" while the Deploy panel says the round is in progress.
|
||||||
|
match: { hw: "rtx6000", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
|
// Flat string, not a predicate: this cell is in-progress with or without
|
||||||
|
// the DSPARK overlay, so there is nothing for the selection to switch on.
|
||||||
|
verificationStatus: "in-progress",
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser minicpm5",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
// The 32GB card is the one where the default KV pool starves decode
|
||||||
|
// CUDA-graph capture: with defaults the pool takes 473,718 tokens / 19 GB
|
||||||
|
// and leaves 4.6 GB, so capture stops around bs=48 and every larger batch
|
||||||
|
// runs eager (3810 tok/s at concurrency 64). The pair below gives back
|
||||||
|
// 4% of the pool -- still hugely oversized for a 2.5B model -- and keeps
|
||||||
|
// batches up to 128 graph-backed (7454 tok/s at the same concurrency).
|
||||||
|
// The two flags go together: raising the cap without freeing the memory
|
||||||
|
// just lets SGLang clamp capture back down.
|
||||||
|
match: { hw: "rtx5090", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
|
verified: true,
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser minicpm5",
|
||||||
|
"--mem-fraction-static 0.75",
|
||||||
|
"--cuda-graph-max-bs 128",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
// GB10 has no discrete VRAM, so `mem_get_info()` reports all 128GB of
|
||||||
|
// unified system memory and the default fraction claims ~89GB for KV --
|
||||||
|
// leaving ~5GB for the OS, which kills the node during warmup with no
|
||||||
|
// traceback and no OOMKilled event. 0.30 is required, not tuning; it
|
||||||
|
// still leaves a 658k-token pool.
|
||||||
|
match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
|
verified: true,
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser minicpm5",
|
||||||
|
"--mem-fraction-static 0.30",
|
||||||
|
"--cuda-graph-max-bs 128",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
};
|
||||||
Reference in New Issue
Block a user