Commit Graph
136 Commits
Author SHA1 Message Date
Thomas Wang f60652a43e [AMD] Update deepseek-v4 PDI and cache policy setting for agentic workload (#39702) 2026-09-15 21:16:30 -07:00
jacky.chengandChangLiu0709 a9bb4d7d45 [AMD] Align Qwen3.5 MI355X HiCache cookbook with kernel / page_first (#39572)
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
2026-09-16 11:22:28 +08:00
amote-i bdf8886ad3 [NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389) 2026-09-15 10:33:14 +08:00
a25f213bc4 [diffusion] docs: give the RTX 5090 its own H3 recipe, measured on a physical desktop (#39373)
Co-authored-by: Mick Qian <mickqian@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 09:15:13 +08:00
ChangLiu0709 242d8a70c0 [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260913, enable TOPK_V2 (#39406) 2026-09-14 21:50:52 +08:00
Theresa Shan 95140a7b0c [docs] DeepSeek-V4: MI355X PD disaggregation recipes for all three strategies (#39396) 2026-09-14 02:11:51 -07:00
jacky.cheng 2f5cc8e33e [AMD] Align Qwen3.5 MI355X cookbook with AttnFP8-V2 and HiCache direct / page_first_direct (#39358) 2026-09-14 12:46:17 +08:00
Thomas Wang ec5fba5777 [AMD] Add dspark config and agentic workload section for deepseek-v4 model (#39252) 2026-09-12 19:32:01 -07:00
Zhang, Jiejing 288627e400 [AMD] Document GLM-5.2 MXFP4 recipe update on MI355X (#39230) 2026-09-12 15:57:15 -07:00
Xinyuan Tong b5a2aebc7e [Docs] GLM-5.3-Flash cookbook: fixed MTP 5/1/6, EP1 + flashinfer_trtllm on Blackwell (#39213) 2026-09-12 12:56:57 -07:00
ChangLiu0709 7bc4eb3740 [AMD] Update MI355X MXFP4 HiCache defaults and quick-reduce quantization for Qwen3.5 cookbook (#39104) 2026-09-12 00:59:32 -07:00
Khoa PhamandClaude Fable 5.1 0d08668821 [Cookbook] Kimi-K3: keep DCP under HiCache L1+L2 with DSPARK (#39190)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 23:14:24 -07:00
ff1ce11348 [diffusion] model: support VDN-H3 with a hybrid_window_attn_h3 backend (#37903)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Haocheng Xi <xihc@berkeley.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-09-12 11:36:32 +08:00
Kevin Mi 5f3606c7b2 [Cookbook][AMD] Kimi-K3 MI350X/MI355X: pin a ROCm image with the DSPARK graph-capture fix, add measured cell numbers (#39029) 2026-09-11 16:10:44 -07:00
df6424967a [docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Jiminator <jimmysh341@gmail.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
2026-09-11 11:58:30 -07:00
Zhang, Jiejing d7c284b894 [AMD] Use the triton DSA backend for GLM-5.2 MXFP4 on MI355X (#39106) 2026-09-11 11:48:50 -07:00
Mohammad Miadh AngkadandMohammad Angkad 52c191da52 [Deps] Retire the CUDA 12 lane (#38404)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-10 16:58:09 -07:00
zijiexiaandClaude Opus 5 4b7331fb77 Make the remaining DeepSeek-V4.1 NVIDIA cells start (#38861)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 11:04:53 -07:00
Yuhao YangandClaude Code a37ded1693 [Cookbook] DeepSeek-V4.1: add the HiCache L2 knob to the Playground (#38844)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-10 18:20:34 +08:00
zijiexiaandClaude Opus 5 5caafd2118 Fix the DeepSeek-V4.1 reasoning example and mark the B300 cells verified (#38839)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 01:49:44 -07:00
Xinyuan Tong 9a2f17f41d Add INT4 and FP4 lanes to the Ling-3.0-flash-VL cookbook (#38527) 2026-09-10 15:23:50 +08:00
zijiexiaandClaude Opus 5 69777c4d36 Add DeepSeek-V4.1 Flash cookbook (#38802)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:15:30 -07:00
Thomas Wang c415f977b8 [AMD] Update v4 args for agentic workload (#38677) 2026-09-09 21:41:05 -07:00
William Hu 0084030179 Add Opt-In for GLM-5.3 Flash breakable prefill CUDA graphs (#38522) 2026-09-09 17:02:08 -07:00
daf66f6670 [Diffusion][SenseNova] support SenseNova-U1.5-8B-MoT (#36606)
Co-authored-by: wuyuefeng <wuyuefeng@noreply.gitcode.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-09-09 10:12:38 +03:00
Cheng Wan db272201a2 [Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement (#38375) 2026-09-08 16:42:12 -07:00
Xinyuan Tong 482e9f257b Add Ling-3.0-flash-VL cookbook (#38434) 2026-09-08 23:00:24 +08:00
Kedar PotdarandPo-Han Huang f4bbf12423 docs(cookbook): Qwen3.5 FP8 on B200/B300 — trtllm-gen MoE + symm mem (#38374)
Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
2026-09-08 09:12:49 +08:00
f4b75b5c36 docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-07 15:33:03 -07:00
zijiexiaandClaude Opus 5 e4008de757 Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 21:30:17 +08:00
Mick ba6d3df69a [diffusion] doc: document verified GB300 and derived GB200 H3 recipes (#38296) 2026-09-07 16:32:29 +08:00
Mick f3d05644db [diffusion] docs+skill: document which components to stream under layerwise offload (#35674) 2026-09-06 23:12:36 +08:00
MickandClaude Fable 5 ade1da017f [diffusion] docs: verify the DGX Spark H3 recipe (#37456)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-06 23:10:27 +08:00
zijiexiaandClaude Opus 5 3b64169f9d [Cookbook] Kimi-K3: add measured B300 1x8 Unified 8k/1k speed numbers (#37878)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:44:00 -07:00
Thomas Wang 225129fe44 [AMD] Update v4 amd cookbook 0903 (#37829) 2026-09-03 22:11:13 -07:00
Jimmy ShongandClaude Fable 5.1 2da5802bfa [Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-03 11:43:01 -07:00
kkandwunhuang dd091f43cd [AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-03 18:44:39 +08:00
Yash Akhauri 02d9b3060a [Docs] Update K2 Horizon MoE model names (#37723) 2026-09-02 23:32:44 -07:00
Yash Akhauri 98ef7d8ae6 docs: add K2 Horizon cookbook recipes and H200 results (#37655) 2026-09-03 11:52:05 +08:00
Xinyuan Tong 3421d4375b [Docs] GLM-5.3-Flash cookbook: drop stale EP caveat, add B300/H100/B200 FP8 speed data (#37576) 2026-09-02 17:16:23 -07:00
Mick f6aed6ec53 [diffusion] doc: rewrite stale diffusion compatibility matrix (#36987) 2026-09-02 23:42:54 +08:00
Liangsheng Yin ebfd8c60e5 [CI] Install sgl-eval from PyPI through the test extra (#37504) 2026-09-02 01:45:25 -07:00
Xiaoyu Zhang 1aa8299d1d [Diffusion] Add cumulative extra-high quality tier (#37422) 2026-09-02 10:26:13 +08:00
zijiexia 6d34a4d3ce [Cookbook] Verify DeepSeek-V4 Flash Vision on GB300 (#37492) 2026-09-01 17:11:41 -07:00
Jimmy ShongandClaude Fable 5.1 ed82bea146 [Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-01 15:42:39 -07:00
zijiexiaandClaude Fable 5 0f18d389b4 [Cookbook] Verify DeepSeek-V4 Flash Vision balanced and high-throughput on B200 (#37468)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 13:09:14 -07:00
Xinyuan Tong 442c7c1e29 [Docs] GLM-5.3-Flash cookbook: add NVFP4 FP8+TRT-LLM benchmark rows (follow-up to #37109) (#37412) 2026-09-01 12:56:22 -07:00
Ankur Singh 3315356cc0 docs(cookbook): enable FlashInfer GDN for Qwen3.5 B200 (#37360) 2026-09-01 11:48:41 -07:00
zijiexiaandClaude Opus 5 dc1ae02684 [Cookbook] Add the DFlash2 speculative option to GLM-5.3 (#37392)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 09:08:28 +00:00
zijiexia 6c72b49a57 Revert "[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)" (#37380) 2026-09-01 01:25:13 -07:00