[Docs] Add NVFP4 quantization to GLM-5.2 cookbook (#29380)

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
zijiexia
2026-06-26 00:40:13 -07:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 5996b54bd3
commit dd56a9f069
8 changed files with 112 additions and 12 deletions
+2 -1
View File
@@ -149,7 +149,8 @@ this table (RTX PRO 6000, GH200, future chips) goes in the model's own `config.h
4. **Fill `cells[]`** with the verified recipes from Phase 1 (replace every EXAMPLE cell;
set `verified: true` only on tested combos), and `modelNames` with real HF slugs,
`dockerImages` for your hw (use the Phase-1 tag, or default `lmsysorg/sglang:dev` — never
a guessed release), `multiNodeHints` only for fabric-specific hw (e.g. gb200).
a guessed release; key by `hw`, or `hw|quant` when one quant on a shared GPU needs its own
image), `multiNodeHints` only for fabric-specific hw (e.g. gb200).
### Site-wiring (do all three)