diff --git a/.claude/skills/cookbook-add-model/SKILL.md b/.claude/skills/cookbook-add-model/SKILL.md index 8aa3e9fd0..85193479f 100644 --- a/.claude/skills/cookbook-add-model/SKILL.md +++ b/.claude/skills/cookbook-add-model/SKILL.md @@ -135,7 +135,8 @@ this table (RTX PRO 6000, GH200, future chips) goes in the model's own `config.h - `supportedHardware` + the EXAMPLE cells: keep your tested families; **delete the `mi*` ids + AMD example cell if no AMD recipe**, etc. A GPU not in the shared catalog (e.g. RTX PRO 6000) → declare it in `config.hardware` and add its id here. - - `playgroundFeatures` axes: delete `megamoe` (non-Blackwell-MoE), `hisparse` + - `playgroundFeatures` axes: remove the `megamoe` backend option + the + `megamoeQuant` block from the `moe` axis (non-Blackwell-MoE), delete `hisparse` (non-DSA), `pdDisagg`/`router` (no PD), the `parsers` axis (no parsers), etc. - `quantizations` / `variants`: drop what the model doesn't ship; collapse `variants` to single `default` if there's no variant axis (then drop the `variant` half of diff --git a/.claude/skills/cookbook-add-model/references/authoring-reference.md b/.claude/skills/cookbook-add-model/references/authoring-reference.md index afc83e491..9d35f2836 100644 --- a/.claude/skills/cookbook-add-model/references/authoring-reference.md +++ b/.claude/skills/cookbook-add-model/references/authoring-reference.md @@ -109,13 +109,12 @@ the `_playground.jsx` header): | Axis key | Widget | Use when | |---|---|---| | `attention` | TP / CP / DP-Attention sub-knobs (DP-Attention is a combined knob: its value is the DP degree AND toggles `--enable-dp-attention`) | Model exposes parallelism knobs in its cells (§2.2) and you want users to override them. | -| `moe` | Backend select + EP knob | Model is MoE and supports multiple `--moe-*-backend` choices. | +| `moe` | Backend select (incl. MegaMoE) + EP knob; picking the MegaMoE backend reveals a Quantization sub-select (W4A8/W4A4) | Model is MoE and supports multiple `--moe-*-backend` choices. For Blackwell MoE kernel-fusion, give the `megamoe` backend option a `requiresHw` (and optional `excludesStrategy`) gate, then add a sibling `megamoeQuant` block (`{stripEnv, options}`): W4A8 = `NUM_MAX` only, W4A4 adds the FP4-activations env vars; both strip the DeepEP dispatch env. | | `parsers` | Multi-toggle | Model has reasoning / tool-call parsers. | | `speculative` | Single-select chip group | Model has spec-decoding presets you want to expose. | | `pdDisagg` | Mode + transfer backend (+ optional per-backend env via `envWhen` hw-gate) + IB device + optional `router{port, command}` | Model supports prefill/decode disaggregation. When a PD role is active and `router` is set, the playground shows the router (SGLang Model Gateway) launch command as a separate companion block and retargets the cURL modal to `router.port` (clients hit the router, not the role servers). | | `hicache` | Enable + storage + write policy | Model is large enough that hierarchical KV cache matters. | | `hisparse` | Enable + host-ratio select; whole card gated on the live PD-Disagg mode being `decode` | DSA-style model (DeepSeek-V3.2 / V4, GLM-5) that supports decode-side hierarchical sparse attention. | -| `megamoe` | Single-select with hw/strategy gating | Blackwell-only kernel fusion variant. | **Per-chip constraints**: any chip entry in any axis can be wrapped with `hide` / `disable` constraint objects: diff --git a/.claude/skills/cookbook-add-model/references/engine-axis.md b/.claude/skills/cookbook-add-model/references/engine-axis.md index 5b67edd60..8d82c656b 100644 --- a/.claude/skills/cookbook-add-model/references/engine-axis.md +++ b/.claude/skills/cookbook-add-model/references/engine-axis.md @@ -1,11 +1,12 @@ # Engine extension: add a new playground feature axis Loaded on demand by the `cookbook-add-model` skill. **Rare** — adding a model -cookbook is data-only and never needs this. The current 8 built-in axes +cookbook is data-only and never needs this. The current 7 built-in axes (`attention`, `moe`, `parsers`, `speculative`, `pdDisagg`, `hicache`, -`hisparse`, `megamoe`) already cover the SGLang feature surface most cookbooks -need. Only add a new axis if a real cookbook needs it and the feature does not -fit any existing axis. Touches `_playground.jsx` only. +`hisparse`) already cover the SGLang feature surface most cookbooks need. +(MegaMoE is not its own axis — it lives inside `moe` as a backend option + +a `megamoeQuant` sub-select.) Only add a new axis if a real cookbook needs it +and the feature does not fit any existing axis. Touches `_playground.jsx` only. For the per-model config/cells/MDX reference see [authoring-reference.md](authoring-reference.md). @@ -94,7 +95,7 @@ Template: // also receives the derived // value (as `derived`) and may use it as a no-op shortcut when the // user's pick matches base. Skip when your axis owns flags that never - // appear in base cells (PD-Disagg / HiCache / MegaMoE). + // appear in base cells (PD-Disagg / HiCache). // deriveFromBase: (cell, fc, h) => ({ ... }) | null, // Optional: hints for the renderer. Currently only pdDisagg uses this @@ -103,7 +104,7 @@ Template: // Returns the axis card JSX. The outer div MUST have key={axisId} so // React can track it in the engine's map loop. Return null for - // axis-level gating (e.g. MegaMoE on Hopper). Lay out as a single + // axis-level gating (e.g. HiSparse when the live PD mode isn't `decode`). Lay out as a single // compact horizontal row: title on the left, fields after. render: ({ axisId, value, setValue, fc, base, s, h, renderChip, renderSelect, derived }) => { if (/* axis-level gating fails */) return null; @@ -147,7 +148,7 @@ Template: the **default** compact control (a `