Add Inkling-Small cookbook (#32951)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai> Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
This commit is contained in:
co-authored by
Zijie Xia
Yanbin Jiang
parent
b61cb5f9de
commit
04edadb34d
@@ -0,0 +1,309 @@
|
||||
---
|
||||
title: Inkling-Small
|
||||
description: "Deploy Inkling-Small with SGLang — launch commands, tuning, and multimodal / reasoning / tool-calling usage for Thinking Machines' Inkling-Small Mixture-of-Experts model."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
## Deployment
|
||||
|
||||
<a id="install" />
|
||||
|
||||
<Accordion title="Install SGLang">
|
||||
|
||||
For all install methods and hardware platforms, see the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
<Tabs>
|
||||
|
||||
<Tab title="Python (pip / uv)">
|
||||
|
||||
Inkling-Small has merged to `main` but isn't in a `pip` release yet — install from source:
|
||||
|
||||
```bash Command
|
||||
pip install --upgrade pip
|
||||
pip install 'git+https://github.com/sgl-project/sglang.git#subdirectory=python'
|
||||
```
|
||||
|
||||
Then run the **Python** output of the command panel below.
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Docker">
|
||||
|
||||
<Note>The Inkling-Small images are being published to [`lmsysorg/sglang`](https://hub.docker.com/r/lmsysorg/sglang/tags) — watch the tag list for status.</Note>
|
||||
|
||||
There are two multi-arch (amd64 / arm64) CUDA builds plus a ROCm build; pick the CUDA build by your CUDA version, not your GPU:
|
||||
|
||||
```bash Command
|
||||
docker pull lmsysorg/sglang:dev-inkling-dspark # CUDA 13
|
||||
docker pull lmsysorg/sglang:dev-cu12-inkling-dspark # CUDA 12
|
||||
docker pull lmsysorg/sglang-rocm:dev-rocm720-mi35x-inkling-dspark # AMD MI350X / MI355X
|
||||
```
|
||||
|
||||
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
|
||||
|
||||
</Tab>
|
||||
|
||||
</Tabs>
|
||||
|
||||
</Accordion>
|
||||
|
||||
Pick your hardware to generate the launch command. Each platform ships a **Balanced** recipe plus **MTP** and **DSpark** (speculative decoding) tiers and a **Long Context (MXFP8 KV)** tier where validated; the **LoRA** variant serves adapters on top of the frozen base model. Set `MAX_LORAS` to the number of distinct adapters you serve (1 is fastest for single-adapter serving).
|
||||
|
||||
import { Deployment } from "/src/snippets/_deployment.jsx";
|
||||
import { config } from "/src/snippets/configs/thinkingmachines/inkling-small.jsx";
|
||||
import { benchmarks } from "/src/snippets/configs/thinkingmachines/inkling-small-benchmarks.jsx";
|
||||
|
||||
<Deployment config={config} benchmarks={benchmarks} />
|
||||
|
||||
<div style={{fontSize: "0.85em", lineHeight: "1.55", color: "#6b7280", margin: "0.5rem 0 1rem 0"}}>
|
||||
<p style={{margin: "0 0 0.3rem 0"}}><strong>Panel controls</strong> (top of the command box):</p>
|
||||
<ul style={{margin: 0, paddingLeft: "1.25rem"}}>
|
||||
<li style={{marginBottom: "0.2rem"}}><strong>⧉ Copy</strong> — copies the current command to your clipboard.</li>
|
||||
<li style={{marginBottom: "0.2rem"}}><strong>$ cURL</strong> — a sample request against <code>localhost:30000</code> to confirm the server is up.</li>
|
||||
<li style={{marginBottom: "0.2rem"}}><strong>⚙ Env</strong> — edits the placeholders (<code>HOST_IP</code>, <code>PORT</code>, <code>NODE_RANK</code>, <code>NODE0_IP</code>) the command and cURL share.</li>
|
||||
<li><strong>Verified / Not Verified</strong> badge — green when the <code>(hw, variant, quant, strategy, nodes)</code> combo has been run end-to-end on real hardware; yellow when auto-derived from a neighbor and not yet re-checked.</li>
|
||||
</ul>
|
||||
</div>
|
||||
|
||||
## Playground
|
||||
|
||||
The Playground is where you experiment with **SGLang features beyond the verified matrix**. The Deploy panel above only emits combinations that have been signed off; the Playground lets you turn on additional knobs on top of whichever cell the Deploy panel is currently showing. The base is read live from your Deploy selection — only your overrides change.
|
||||
|
||||
Lines highlighted **green** are added by your overrides; lines with **red strikethrough** were in the verified base but stripped by an override. Any change flips the badge to **Not Verified** until the new configuration is run end-to-end.
|
||||
|
||||
import { Playground } from "/src/snippets/_playground.jsx";
|
||||
|
||||
<Playground config={config} />
|
||||
|
||||
## 1. Model Introduction
|
||||
|
||||
**Inkling-Small** is a Mixture-of-Experts model from Thinking Machines with **open weights** (BF16 and NVFP4 checkpoints below), in the same architecture family as Inkling. It handles text, image, and audio inputs natively, and exposes a **variable reasoning-effort** control to trade latency and cost against answer quality. This page covers serving Inkling-Small on SGLang, including its **MTP** speculative-decoding path and long-context prefix caching (unified radix cache + HiCache).
|
||||
|
||||
**Resources:** HuggingFace — [Inkling-Small](https://huggingface.co/thinkingmachines/Inkling-Small) (BF16) · [Inkling-Small-NVFP4](https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4).
|
||||
|
||||
## 2. Configuration Tips
|
||||
|
||||
**Multimodal.** The recipes pass `--enable-multimodal` so the server accepts image and audio inputs alongside text — drop it for text-only serving.
|
||||
|
||||
**Memory pool ratios.** `--swa-full-tokens-ratio` and `--mamba-full-memory-ratio` (both default `0.1`) size the SWA and Mamba/sconv state pools; tune them to your workload's usage.
|
||||
|
||||
**MTP needs `--enable-multi-layer-eagle`.** The MTP recipe drives Inkling-Small's multi-layer draft head; without this flag the standard EAGLE worker runs against it and outputs garbage.
|
||||
|
||||
**Reasoning effort.** Pass `reasoning_effort` as one of the named levels below; requests that omit it default to `high`, and `max` is the strongest. Each level maps to an internal effort value (max at `0.99`):
|
||||
|
||||
<table style={{width: "60%", borderCollapse: "collapse"}}>
|
||||
<thead>
|
||||
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||
<th style={{textAlign: "left", padding: "8px 12px", fontWeight: 700}}>reasoning_effort</th>
|
||||
<th style={{textAlign: "left", padding: "8px 12px", fontWeight: 700}}>value</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>none</code></td><td style={{padding: "6px 12px"}}>0.0</td></tr>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>minimal</code></td><td style={{padding: "6px 12px"}}>0.1</td></tr>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>low</code></td><td style={{padding: "6px 12px"}}>0.2</td></tr>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>medium</code></td><td style={{padding: "6px 12px"}}>0.7</td></tr>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>high</code></td><td style={{padding: "6px 12px"}}>0.9</td></tr>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>xhigh</code></td><td style={{padding: "6px 12px"}}>0.99</td></tr>
|
||||
<tr><td style={{padding: "6px 12px"}}><code>max</code></td><td style={{padding: "6px 12px"}}>0.99</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
## 3. Advanced Usage
|
||||
|
||||
### 3.1 Reasoning
|
||||
|
||||
Enable the `inkling` reasoning parser (toggle **Reasoning Parser** in the **Parsers** card of the [Playground above](#playground)) to separate thinking from the final answer into `reasoning_content` vs `content`.
|
||||
|
||||
<Accordion title="Reasoning Example (Python)">
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
|
||||
|
||||
resp = client.chat.completions.create(
|
||||
model="thinkingmachines/Inkling-Small-NVFP4",
|
||||
messages=[{"role": "user", "content": "What is 17 times 24?"}],
|
||||
extra_body={"chat_template_kwargs": {"thinking": True}},
|
||||
)
|
||||
msg = resp.choices[0].message
|
||||
print("Reasoning:", getattr(msg, "reasoning_content", None))
|
||||
print("Answer:", msg.content)
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Example Output">
|
||||
|
||||
```text Output
|
||||
Reasoning: The user is asking for the product of 17 and 24. Let me calculate that.
|
||||
|
||||
17 × 24
|
||||
|
||||
I can break this down:
|
||||
17 × 20 = 340
|
||||
17 × 4 = 68
|
||||
340 + 68 = 408
|
||||
|
||||
Alternatively:
|
||||
24 × 10 = 240
|
||||
24 × 7 = 168
|
||||
240 + 168 = 408
|
||||
|
||||
So the answer is 408.
|
||||
Answer: 17 times 24 is **408**.
|
||||
|
||||
Here's a quick breakdown:
|
||||
- 17 × 20 = 340
|
||||
- 17 × 4 = 68
|
||||
- 340 + 68 = **408**
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
### 3.2 Tool Calling
|
||||
|
||||
Enable the `inkling` tool-call parser (toggle **Tool Call Parser** in the **Parsers** card of the [Playground above](#playground)) to surface structured tool calls via `message.tool_calls`.
|
||||
|
||||
<Accordion title="Tool Calling Example (Python)">
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
|
||||
|
||||
tools = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get the current weather for a location",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {"location": {"type": "string", "description": "The city name"}},
|
||||
"required": ["location"],
|
||||
},
|
||||
},
|
||||
}
|
||||
]
|
||||
|
||||
resp = client.chat.completions.create(
|
||||
model="thinkingmachines/Inkling-Small-NVFP4",
|
||||
messages=[{"role": "user", "content": "What's the weather in Beijing?"}],
|
||||
tools=tools,
|
||||
)
|
||||
msg = resp.choices[0].message
|
||||
print("Reasoning:", getattr(msg, "reasoning_content", None))
|
||||
print("Content:", msg.content)
|
||||
print("Tool calls:", msg.tool_calls)
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Example Output">
|
||||
|
||||
```text Output
|
||||
Reasoning: The user is asking for the weather in Beijing. I have a tool called `get_weather` that can get the current weather for a location. Let me call it with "Beijing" as the location.
|
||||
Content:
|
||||
Tool calls: [ChatCompletionMessageFunctionToolCall(id='call_98f772f3a0044f45b80c5ba5', function=Function(arguments='{"location": "Beijing"}', name='get_weather'), type='function', index=0)]
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
### 3.3 Multimodal Input (Image + Audio)
|
||||
|
||||
Inkling-Small is multimodal: a single user message can mix **text**, **images**, and **audio**. Pass each media item as its own content part — `image_url` for images, `audio_url` for audio — with the `url` set to either an HTTP(S) link or a base64 `data:` URI. The server must be started with `--enable-multimodal` (already included in every recipe above).
|
||||
|
||||
<Accordion title="Image + Audio Example (Python)">
|
||||
|
||||
```python Example
|
||||
import base64
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
|
||||
|
||||
with open("image.png", "rb") as f:
|
||||
image_b64 = base64.b64encode(f.read()).decode()
|
||||
with open("audio.wav", "rb") as f:
|
||||
audio_b64 = base64.b64encode(f.read()).decode()
|
||||
|
||||
resp = client.chat.completions.create(
|
||||
model="thinkingmachines/Inkling-Small-NVFP4",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image_b64}"}},
|
||||
{"type": "audio_url", "audio_url": {"url": f"data:audio/wav;base64,{audio_b64}"}},
|
||||
{"type": "text", "text": "Describe the image, then transcribe the audio."},
|
||||
],
|
||||
}
|
||||
],
|
||||
max_tokens=1024,
|
||||
)
|
||||
print(resp.choices[0].message.content)
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Note>
|
||||
Images and audio can be sent as public HTTP(S) URLs instead of base64 — e.g. `{"type": "image_url", "image_url": {"url": "https://.../photo.jpg"}}`. Use one content part per media item; mix as many as the context budget allows.
|
||||
</Note>
|
||||
|
||||
### 3.4 LoRA (Serving Adapters)
|
||||
|
||||
The **LoRA** deploy variant serves adapters on top of the frozen base model. Its launch command adds `--enable-lora --lora-paths lora0={{ADAPTER_PATH}} --max-loras-per-batch {{MAX_LORAS}}` — each adapter is registered under the **name** to the left of `=` (here `lora0`). Adapters can also be added/removed at runtime via the `POST /load_lora_adapter` endpoint. To serve several adapters, pass multiple `--lora-paths name=path` at launch and reference each by its name.
|
||||
|
||||
Pick the adapter per request by that name — either in the `model` field with `base-model:adapter` syntax (recommended), or explicitly via `lora_path` in `extra_body`:
|
||||
|
||||
<Accordion title="LoRA Example (Python)">
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
|
||||
|
||||
# Option A (recommended): "<model>:<adapter-name>" in the model field
|
||||
resp = client.chat.completions.create(
|
||||
model="thinkingmachines/Inkling-Small-NVFP4:lora0",
|
||||
messages=[{"role": "user", "content": "Summarize the changelog."}],
|
||||
)
|
||||
|
||||
# Option B: explicit lora_path via extra_body
|
||||
resp = client.chat.completions.create(
|
||||
model="thinkingmachines/Inkling-Small-NVFP4",
|
||||
messages=[{"role": "user", "content": "Summarize the changelog."}],
|
||||
extra_body={"lora_path": "lora0"},
|
||||
)
|
||||
|
||||
print(resp.choices[0].message.content)
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Note>
|
||||
One adapter per request — omit the `:adapter` suffix (and `lora_path`) to hit the base model. Different requests **in the same batch** may use different adapters; the number of *distinct* adapters co-resident in a batch is capped by `--max-loras-per-batch` (the `MAX_LORAS` field, default `1`). If both `model:adapter` and `lora_path` are supplied, the `model` suffix takes precedence.
|
||||
</Note>
|
||||
|
||||
### 3.5 HiCache (Hierarchical KV Caching)
|
||||
|
||||
Inkling-Small serves on SGLang's **unified radix cache**: the historically separate full-attention, SWA, and Mamba/sconv caches are combined into one radix tree with typed components, and native HiCache offloads cold prefix pages across tiers (GPU HBM → host DRAM → disk / remote). This expands effective prefix-cache capacity for multi-turn and long-context workloads.
|
||||
|
||||
To enable HiCache, open the **HiCache** card in the [Playground above](#playground) and flip **Enable**, then pick a storage backend (`file` / `mooncake` / `nixl`) for the L3 tier. The Write policy defaults to `write_through`.
|
||||
|
||||
### 3.6 Long Context (MXFP8 KV)
|
||||
|
||||
The **Long Context** deploy strategy adds `--kv-cache-dtype mxfp8` on top of the Balanced recipe. KV entries are stored as block-scaled MXFP8 instead of BF16, so the SWA + Mamba/sconv memory pool holds roughly 2x as many tokens on the same GPU. Use it when you're context-bound or concurrency-bound.
|
||||
|
||||
**Blackwell only.** MXFP8 KV cache requires Blackwell (B200 / B300 / GB200 / GB300), it's not offered on Hopper (H200).
|
||||
|
||||
The tradeoff is not just a ~5% decode latency penalty from the extra quantize/dequantize work versus BF16 KV — storing KV in MXFP8 also introduces some accuracy loss at long context lengths. Treat it as a capacity lever, not a speed one — stay on **Balanced** if you have headroom in the memory pool and just want lower latency or maximum output quality.
|
||||
|
||||
To try it, select the **Long Context** strategy in the Deploy panel above for any NVFP4 cell; the panel regenerates the launch command with `--kv-cache-dtype mxfp8` inserted. Verified end-to-end on B200, B300, and GB300.
|
||||
|
||||
### 3.7 DSpark (Speculative Decoding)
|
||||
|
||||
The **DSpark** deploy strategy is the second speculative-decoding path for Inkling-Small. Unlike **MTP**, which drives Inkling-Small's own multi-layer draft head, DSpark runs a **separate draft checkpoint** — `RadixArk/Inkling-Small-DSpark-Preview` — served unquantized alongside the NVFP4 target.
|
||||
|
||||
DSpark support ships in the images listed in §1 (`dev-inkling-dspark` for CUDA 13, `dev-cu12-inkling-dspark` for CUDA 12), so no separate build is needed. Verified end-to-end on B200 (TP=8, NVFP4).
|
||||
@@ -34,10 +34,9 @@ Then run the **Python** output of the command panel below.
|
||||
There are two multi-arch (amd64 / arm64) CUDA builds plus a ROCm build; pick the CUDA build by your CUDA version, not your GPU:
|
||||
|
||||
```bash Command
|
||||
docker pull lmsysorg/sglang:inkling-cu13 # CUDA 13
|
||||
docker pull lmsysorg/sglang:inkling-cu12 # CUDA 12
|
||||
docker pull lmsysorg/sglang:inkling-rocm700-mi35x # AMD MI350X / MI355X
|
||||
docker pull lmsysorg/sglang:dev-cu13-inkling-dspark # CUDA 13 + DSpark support
|
||||
docker pull lmsysorg/sglang:dev-inkling-dspark # CUDA 13
|
||||
docker pull lmsysorg/sglang:dev-cu12-inkling-dspark # CUDA 12
|
||||
docker pull lmsysorg/sglang-rocm:dev-rocm720-mi35x-inkling-dspark # AMD MI350X / MI355X
|
||||
```
|
||||
|
||||
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
|
||||
@@ -307,4 +306,4 @@ To try it, select the **Long Context** strategy in the Deploy panel above for an
|
||||
|
||||
The **DSpark** deploy strategy is the second speculative-decoding path for Inkling. Unlike **MTP**, which drives Inkling's own multi-layer draft head, DSpark runs a **separate draft checkpoint** — `RadixArk/Inkling-DSpark-Preview` — served unquantized alongside the NVFP4 target.
|
||||
|
||||
It needs an SGLang build with DSpark support — `lmsysorg/sglang:dev-cu13-inkling-dspark`, which the Deploy panel's **Docker** command uses for this tier; the `inkling-cu13` / `inkling-cu12` images don't carry DSpark yet. Verified end-to-end on B200 (TP=8, NVFP4).
|
||||
DSpark support ships in the images listed in §1 (`dev-inkling-dspark` for CUDA 13, `dev-cu12-inkling-dspark` for CUDA 12), so no separate build is needed. Verified end-to-end on B200 (TP=8, NVFP4).
|
||||
|
||||
+2
-1
@@ -987,7 +987,8 @@
|
||||
{
|
||||
"group": "Thinking Machines",
|
||||
"pages": [
|
||||
"cookbook/autoregressive/ThinkingMachines/Inkling"
|
||||
"cookbook/autoregressive/ThinkingMachines/Inkling",
|
||||
"cookbook/autoregressive/ThinkingMachines/Inkling-Small"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
@@ -163,8 +163,8 @@ export const Deployment = ({ config, benchmarks }) => {
|
||||
color: isDark ? "#e5e7eb" : "#374151",
|
||||
whiteSpace: "pre-wrap", overflowX: "auto", margin: 0,
|
||||
},
|
||||
// Amber callout under the command when speculative decoding (MTP) is on
|
||||
// but --max-running-requests isn't set (SGLang then caps it at 48).
|
||||
// Amber callout under the command when speculative decoding (MTP, DSpark, ...)
|
||||
// is on but --max-running-requests isn't set (SGLang then caps it at 48).
|
||||
mtpWarn: {
|
||||
margin: "8px 0 0", padding: "8px 12px", borderRadius: "8px",
|
||||
fontSize: "12px", lineHeight: "1.45",
|
||||
@@ -1176,13 +1176,37 @@ export const Deployment = ({ config, benchmarks }) => {
|
||||
return { ...cell, flags };
|
||||
})();
|
||||
const command = renderCommand(cellWithRatio, sel, env, runMode);
|
||||
// MTP hint on the EFFECTIVE flags — speculation arrives via the Spec Decode
|
||||
// overlay, never the cell. SGLang resets --max-running-requests to 48 when
|
||||
// spec is on and it's unset.
|
||||
// Speculative-decoding hint on the EFFECTIVE flags — speculation can arrive via
|
||||
// the Spec Decode overlay as well as the cell. SGLang resets
|
||||
// --max-running-requests to 48 when spec is on and it's unset; verified for both
|
||||
// EAGLE/MTP and DSPARK (server_args reports max_running_requests=48 either way).
|
||||
const effFlags = cell ? [...overlayStrip(cell.flags, sel), ...overlayFlags(sel)] : [];
|
||||
const mtpHint =
|
||||
effFlags.some((f) => f.split(/[\s=]/)[0] === "--speculative-algorithm") &&
|
||||
!effFlags.some((f) => f.split(/[\s=]/)[0] === "--max-running-requests");
|
||||
const specAlgoFlag = effFlags.find(
|
||||
(f) => f.split(/[\s=]/)[0] === "--speculative-algorithm");
|
||||
const specMrrFlag = effFlags.find(
|
||||
(f) => f.split(/[\s=]/)[0] === "--max-running-requests");
|
||||
// Two cases, both worth surfacing when speculation is on:
|
||||
// mtpHint — the flag is MISSING, so SGLang silently caps at 48 (a hazard)
|
||||
// specPinnedHint— the recipe PINS it, which is safe but is a fixed number the
|
||||
// reader still has to match to their own concurrency
|
||||
const mtpHint = !!specAlgoFlag && !specMrrFlag;
|
||||
const specPinnedHint = !!specAlgoFlag && !!specMrrFlag;
|
||||
const specMrrValue = specMrrFlag
|
||||
? (specMrrFlag.split(/[\s=]/).filter(Boolean)[1] || "")
|
||||
: "";
|
||||
// Name the algorithm in the banner rather than hardcoding "MTP" — the same reset
|
||||
// applies to DSpark and friends, and a DSpark user reading "(MTP)" would be
|
||||
// misled. The cookbook calls the EAGLE-based path MTP, so keep that mapping.
|
||||
const SPEC_ALGO_LABEL = {
|
||||
EAGLE: "MTP", EAGLE3: "MTP", FROZEN_KV_MTP: "MTP",
|
||||
DSPARK: "DSpark", DFLASH: "DFlash", NGRAM: "N-gram",
|
||||
STANDALONE: "standalone draft",
|
||||
};
|
||||
const specAlgoName = (() => {
|
||||
if (!specAlgoFlag) return "MTP";
|
||||
const v = specAlgoFlag.split(/[\s=]/).filter(Boolean)[1] || "";
|
||||
return SPEC_ALGO_LABEL[v.toUpperCase()] || v || "MTP";
|
||||
})();
|
||||
// cell.warn may embed [label](#anchor) links — rendered as scrollIntoView
|
||||
// buttons, not hrefs, so the hash (which carries the selection) isn't overwritten.
|
||||
const renderWarn = (text) => {
|
||||
@@ -1413,7 +1437,12 @@ export const Deployment = ({ config, benchmarks }) => {
|
||||
{cell && cell.warn && <div style={s.mtpWarn}>⚠️ {renderWarn(cell.warn)}</div>}
|
||||
{mtpHint && (
|
||||
<div style={s.mtpWarn}>
|
||||
⚠️ Speculative decoding (MTP) is on — SGLang resets <code>--max-running-requests</code> to <strong>48</strong> when it isn't set. Add <code>--max-running-requests <N></code> sized for your target concurrency.
|
||||
⚠️ Speculative decoding ({specAlgoName}) is on — SGLang resets <code>--max-running-requests</code> to <strong>48</strong> when it isn't set. Add <code>--max-running-requests <N></code> sized for your target concurrency.
|
||||
</div>
|
||||
)}
|
||||
{specPinnedHint && (
|
||||
<div style={s.mtpWarn}>
|
||||
ℹ️ Speculative decoding ({specAlgoName}) is on and this recipe pins <code>--max-running-requests</code> to <strong>{specMrrValue}</strong>. Adjust it to match your target concurrency — if you remove the flag, SGLang falls back to <strong>48</strong>.
|
||||
</div>
|
||||
)}
|
||||
</>)}
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
// One entry per cell `match` tuple. `accuracy` is keyed to
|
||||
// config.accuracyLabels in inkling-small.jsx.
|
||||
|
||||
export const benchmarks = [
|
||||
{ match: { hw: "b200" , variant: "default" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" },
|
||||
sglang_version: "dev-inkling-dspark (b7252cc)",
|
||||
accuracy: { aime26_pct: 95.42, bfcl_pct: 76.54, mmau_pct: 76.30 } },
|
||||
{ match: { hw: "b300" , variant: "default" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" },
|
||||
sglang_version: "dev (cb12a15)",
|
||||
accuracy: { gsm8k_pct: 96.29 } },
|
||||
{ match: { hw: "gb200" , variant: "default" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "gb300" , variant: "default" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" },
|
||||
sglang_version: "dev (cb12a15)",
|
||||
accuracy: { gsm8k_pct: 96.66 } },
|
||||
{ match: { hw: "h200" , variant: "default" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" },
|
||||
sglang_version: "dev-inkling-dspark (b7252cc)",
|
||||
accuracy: { aime26_pct: 95.00, bfcl_pct: 76.02, mmau_pct: 74.70 } },
|
||||
{ match: { hw: "mi350x" , variant: "default" , quant: "bf16" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "mi355x" , variant: "default" , quant: "bf16" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "b200" , variant: "default" , quant: "nvfp4" , strategy: "mtp" , nodes: "single" },
|
||||
sglang_version: "dev-inkling-dspark (b7252cc)",
|
||||
accuracy: { aime26_pct: 96.25, bfcl_pct: 77.57, mmau_pct: 77.20 } },
|
||||
{ match: { hw: "b300" , variant: "default" , quant: "nvfp4" , strategy: "mtp" , nodes: "single" },
|
||||
sglang_version: "dev-cu13-inkling-dspark (86ccfef)",
|
||||
accuracy: { gsm8k_pct: 96.06 } },
|
||||
{ match: { hw: "gb200" , variant: "default" , quant: "nvfp4" , strategy: "mtp" , nodes: "single" } },
|
||||
{ match: { hw: "gb300" , variant: "default" , quant: "nvfp4" , strategy: "mtp" , nodes: "single" },
|
||||
sglang_version: "dev-cu13-inkling-dspark (86ccfef)",
|
||||
accuracy: { gsm8k_pct: 96.44 } },
|
||||
{ match: { hw: "h200" , variant: "default" , quant: "nvfp4" , strategy: "mtp" , nodes: "single" },
|
||||
sglang_version: "dev-inkling-dspark (b7252cc)",
|
||||
accuracy: { aime26_pct: 95.83, bfcl_pct: 76.09, mmau_pct: 76.80 } },
|
||||
{ match: { hw: "b200" , variant: "default" , quant: "nvfp4" , strategy: "dspark" , nodes: "single" },
|
||||
sglang_version: "dev-inkling-dspark (b7252cc)",
|
||||
accuracy: { aime26_pct: 95.83, bfcl_pct: 76.31, mmau_pct: 76.80 } },
|
||||
{ match: { hw: "b300" , variant: "default" , quant: "nvfp4" , strategy: "dspark" , nodes: "single" },
|
||||
sglang_version: "dev-cu13-inkling-dspark (86ccfef)",
|
||||
accuracy: { gsm8k_pct: 96.21 } },
|
||||
{ match: { hw: "gb300" , variant: "default" , quant: "nvfp4" , strategy: "dspark" , nodes: "single" },
|
||||
sglang_version: "dev-cu13-inkling-dspark (86ccfef)",
|
||||
accuracy: { gsm8k_pct: 95.83 } },
|
||||
{ match: { hw: "h200" , variant: "default" , quant: "nvfp4" , strategy: "dspark" , nodes: "single" },
|
||||
sglang_version: "dev-inkling-dspark (b7252cc)",
|
||||
accuracy: { aime26_pct: 96.25, bfcl_pct: 76.68, mmau_pct: 76.50 } },
|
||||
{ match: { hw: "b200" , variant: "default" , quant: "nvfp4" , strategy: "long_context" , nodes: "single" },
|
||||
sglang_version: "dev (8fbf960)",
|
||||
accuracy: { gsm8k_pct: 96.13 } },
|
||||
{ match: { hw: "b300" , variant: "default" , quant: "nvfp4" , strategy: "long_context" , nodes: "single" },
|
||||
sglang_version: "dev (cb12a15)",
|
||||
accuracy: { gsm8k_pct: 95.91 } },
|
||||
{ match: { hw: "gb200" , variant: "default" , quant: "nvfp4" , strategy: "long_context" , nodes: "single" } },
|
||||
{ match: { hw: "gb300" , variant: "default" , quant: "nvfp4" , strategy: "long_context" , nodes: "single" },
|
||||
sglang_version: "dev (cb12a15)",
|
||||
accuracy: { gsm8k_pct: 96.21 } },
|
||||
{ match: { hw: "gb300" , variant: "default" , quant: "bf16" , strategy: "balanced" , nodes: "multi-2" } },
|
||||
{ match: { hw: "gb300" , variant: "default" , quant: "bf16" , strategy: "mtp" , nodes: "multi-2" } },
|
||||
{ match: { hw: "b300" , variant: "default" , quant: "bf16" , strategy: "balanced" , nodes: "single" },
|
||||
sglang_version: "dev (cb12a15)",
|
||||
accuracy: { gsm8k_pct: 96.29 } },
|
||||
{ match: { hw: "b300" , variant: "default" , quant: "bf16" , strategy: "mtp" , nodes: "single" },
|
||||
sglang_version: "dev-cu13-inkling-dspark (86ccfef)",
|
||||
accuracy: { gsm8k_pct: 96.36 } },
|
||||
{ match: { hw: "b200" , variant: "default" , quant: "bf16" , strategy: "balanced" , nodes: "multi-2" } },
|
||||
{ match: { hw: "b200" , variant: "default" , quant: "bf16" , strategy: "mtp" , nodes: "multi-2" } },
|
||||
{ match: { hw: "b200" , variant: "lora" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "b300" , variant: "lora" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "gb200" , variant: "lora" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "gb300" , variant: "lora" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "h200" , variant: "lora" , quant: "nvfp4" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "gb300" , variant: "lora" , quant: "bf16" , strategy: "balanced" , nodes: "single" } },
|
||||
{ match: { hw: "h200" , variant: "lora" , quant: "bf16" , strategy: "balanced" , nodes: "single" } },
|
||||
];
|
||||
File diff suppressed because it is too large
Load Diff
@@ -72,19 +72,18 @@ export const config = {
|
||||
-H 'Content-Type: application/json' \\
|
||||
-d '{ "model": "{{MODEL_NAME}}", "messages": [{"role":"user","content":"Hello"}] }'`,
|
||||
|
||||
// NVIDIA: two multi-arch CUDA builds (inkling-cu12 / inkling-cu13) — pick by your
|
||||
// CUDA version, not by GPU. AMD: inkling-rocm700-mi35x. Panel defaults to cu13.
|
||||
// The DSpark tier needs its own preview build (DSpark isn't in the inkling-cu1x
|
||||
// images yet), so it takes a `hw|quant|strategy` key.
|
||||
// NVIDIA: two multi-arch CUDA builds (dev-inkling-dspark for CUDA 13,
|
||||
// dev-cu12-inkling-dspark for CUDA 12) — pick by your CUDA version, not by GPU.
|
||||
// Panel defaults to cu13. AMD: dev-rocm720-mi35x-inkling-dspark (sglang-rocm repo).
|
||||
// All tiers ship from the same images, DSpark included.
|
||||
dockerImages: {
|
||||
"b200|nvfp4|dspark": "lmsysorg/sglang:dev-cu13-inkling-dspark",
|
||||
h200: "lmsysorg/sglang:inkling-cu13",
|
||||
b200: "lmsysorg/sglang:inkling-cu13",
|
||||
b300: "lmsysorg/sglang:inkling-cu13",
|
||||
gb200: "lmsysorg/sglang:inkling-cu13",
|
||||
gb300: "lmsysorg/sglang:inkling-cu13",
|
||||
mi350x: "lmsysorg/sglang:inkling-rocm700-mi35x",
|
||||
mi355x: "lmsysorg/sglang:inkling-rocm700-mi35x",
|
||||
h200: "lmsysorg/sglang:dev-inkling-dspark",
|
||||
b200: "lmsysorg/sglang:dev-inkling-dspark",
|
||||
b300: "lmsysorg/sglang:dev-inkling-dspark",
|
||||
gb200: "lmsysorg/sglang:dev-inkling-dspark",
|
||||
gb300: "lmsysorg/sglang:dev-inkling-dspark",
|
||||
mi350x: "lmsysorg/sglang-rocm:dev-rocm720-mi35x-inkling-dspark",
|
||||
mi355x: "lmsysorg/sglang-rocm:dev-rocm720-mi35x-inkling-dspark",
|
||||
},
|
||||
|
||||
github: {
|
||||
@@ -689,9 +688,8 @@ export const config = {
|
||||
// ====================================================================
|
||||
// DSpark (speculative decoding) — separate draft checkpoint
|
||||
// (RadixArk/Inkling-DSpark-Preview, served unquantized) instead of Inkling's
|
||||
// own MTP head, so it needs a build carrying DSpark support: the
|
||||
// `dev-cu13-inkling-dspark` image (see the dockerImages key above).
|
||||
// The draft weights sit outside the FP4 target, hence mem-fraction 0.68.
|
||||
// own MTP head. The draft weights sit outside the FP4 target, hence
|
||||
// mem-fraction 0.68.
|
||||
// B200 verified end-to-end.
|
||||
// ====================================================================
|
||||
{
|
||||
@@ -714,6 +712,7 @@ export const config = {
|
||||
"--mem-fraction-static 0.68",
|
||||
"--swa-full-tokens-ratio 0.1",
|
||||
"--mamba-full-memory-ratio 0.1",
|
||||
"--enable-multimodal",
|
||||
"--max-running-requests 68",
|
||||
"--reasoning-parser inkling",
|
||||
"--tool-call-parser inkling",
|
||||
|
||||
Reference in New Issue
Block a user