Commit Graph
151 Commits
Author SHA1 Message Date
Jimmy ShongandClaude Fable 5.1 2da5802bfa [Cookbook] DeepSeek-V4 DGX Spark: v2 image + Flash Official NVFP4 and Flash Vision FP4 cells (#37737)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-03 11:43:01 -07:00
triple-muandmickqian bf71035d39 [diffusion] MiniMax-H3: tiered AdaLN plan cache (pinned-host tier + per-plan LRU) (#37266)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
2026-09-03 22:09:21 +08:00
kkandwunhuang dd091f43cd [AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-03 18:44:39 +08:00
Yash Akhauri 02d9b3060a [Docs] Update K2 Horizon MoE model names (#37723) 2026-09-02 23:32:44 -07:00
Yash Akhauri 98ef7d8ae6 docs: add K2 Horizon cookbook recipes and H200 results (#37655) 2026-09-03 11:52:05 +08:00
Faradawn Yang 9c70d22721 Update GLM-5.2 NVFP4 B200/B300 for AgentX HiCache (#35368) 2026-09-02 14:39:09 -07:00
Mick f6aed6ec53 [diffusion] doc: rewrite stale diffusion compatibility matrix (#36987) 2026-09-02 23:42:54 +08:00
Kevin MiandClaude Fable 5 f586654518 [diffusion] feat: support FastH3 (4-step VSA-distilled MiniMax-H3) with a VSA-H3 attention backend (#37480)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-02 21:39:54 +08:00
Liangsheng Yin ebfd8c60e5 [CI] Install sgl-eval from PyPI through the test extra (#37504) 2026-09-02 01:45:25 -07:00
4b329482e8 [diffusion] feat: support cube sparse attention for minimax h3 (#34893)
Co-authored-by: zhenaozhenfu <zhenaozhenfu@minimaxi.com>
Co-authored-by: Reynor <reynor@minimaxi.com>
2026-09-02 15:21:33 +08:00
Xiaoyu Zhang 1aa8299d1d [Diffusion] Add cumulative extra-high quality tier (#37422) 2026-09-02 10:26:13 +08:00
Jimmy ShongandClaude Fable 5.1 ed82bea146 [Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-01 15:42:39 -07:00
zijiexiaandClaude Opus 5 dc1ae02684 [Cookbook] Add the DFlash2 speculative option to GLM-5.3 (#37392)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 09:08:28 +00:00
zijiexia 6c72b49a57 Revert "[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)" (#37380) 2026-09-01 01:25:13 -07:00
zijiexiaandClaude Fable 5 379e33d87e [Cookbook] Add NVFP4 options for DeepSeek-V4 Flash Official (0731) and Pro Official (0813) (#37351)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 08:03:39 +00:00
Xinyuan Tong 60548501bb [Docs] Add NVFP4 section to GLM-5.3-Flash cookbook (#37109) 2026-09-01 14:14:30 +08:00
Yuan Luoandluoyuan.luo 5b04408784 [MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE (#34967)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-08-31 20:04:41 -07:00
71cee04ebe [Diffusion] Optimize Qwen-Image TP collectives and attention (#36680)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:28:37 +08:00
zijiexia 455232de6e [Cookbook] Enable DSpark on the DeepSeek-V4 Flash Vision low-latency recipes (#37301) 2026-08-31 16:43:54 -07:00
zijiexia 88cf5c9541 [Cookbook] Add DeepSeek-V4-Flash-Vision-Exp to the DeepSeek-V4 page (#37293) 2026-08-31 14:52:28 -07:00
Mohammad Miadh Angkad 4f761e8649 [Deps] Bump FlashInfer to 0.6.18 (#36954) 2026-08-30 19:02:39 -07:00
WenhaoZhang e9a7157615 [diffusion] feat: allow cache-dit with dit layerwise offload (#35858) 2026-08-30 20:55:41 +08:00
Thomas Wang 7399c2b558 [AMD] Update v4 amd cookbook 0830 (#37092) 2026-08-29 23:55:28 -07:00
Xiaoyu Zhang 0e1146d04f [diffusion] optimization: optimize Pi0.5 inference and bounded graph serving (#34599) 2026-08-29 14:47:32 +08:00
Liangsheng Yin a25df83fe3 [Cookbook] Run accuracy benchmarks through sgl-eval (#36977) 2026-08-28 23:29:38 -07:00
Thomas Wang 89816a21a1 [AMD] Update v4 amd cookbook 0828 (#36828) 2026-08-28 17:20:22 -07:00
Xiaoyu Zhang 50bc1a3767 [diffusion] Keep Cosmos3 Nano resident on 96 GB GPUs (#36641) 2026-08-29 07:40:46 +08:00
395c2258c3 [Docs] Add GLM-5.3 cookbook (#36827)
Co-authored-by: JustinTong0323 <xinyuantong.cs@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Mohammad Angkad <176301910+mmangkad@users.noreply.github.com>
2026-08-28 14:56:06 +00:00
Mick 803b4fb31c [diffusion] refactor: scope model-specific API parameters (#35613) 2026-08-28 19:08:30 +08:00
Haoguang Cai 989e51ba9c [Docs] Rename Tencent cookbook page titles to "Hy4 preview" / "Hy3 preview" (#36823) 2026-08-28 01:31:52 -07:00
zijiexiaandClaude Fable 5 2960d69622 [Cookbook] Hy4-Preview follow-ups: runtime-accurate recipes + released-model info (#36808)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 00:19:07 -07:00
zijiexiaandClaude Fable 5 1948b61ad4 [Cookbook] Add the Hy4-Preview model page (Tencent) (#36804)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 23:26:13 -07:00
zijiexiaandClaude Fable 5 43b5a57dbb [Docs] Feature GLM-5.3-Flash in the popular-models banner (#36784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 21:56:10 -07:00
zijiexiaandClaude Opus 5 6ccfeb59bc cookbook: add a Speculative card to the GLM-5.3-Flash playground (#36740)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 16:03:59 -07:00
Xinyuan Tong d1f14431fd GLM-5.3-Flash cookbook: HiCache for LL, fusion-flag drop, EAGLE, default-cell numbers, DCP4 overlay (#36544) 2026-08-28 03:18:25 +08:00
zijiexiaandClaude Opus 5 46a544e0a0 [Docs] GLM-5.3-Flash: point at compute-mamba-ratio for the KDA/KV pool split (#36719)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 12:16:35 -07:00
zijiexia 636a6f7dba cookbook: fix GLM-5.3-Flash speculative flag, size Hopper memory, record GSM8K (#36660) 2026-08-27 02:19:47 -07:00
andyluo7 0f7b5b8b2a [AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608) 2026-08-27 05:19:57 +00:00
Xinyuan Tong e27a7fac77 GLM-5.3-Flash cookbook: default Blackwell recipes to FP8 KV + TRT-LLM DSA (#36519) 2026-08-27 00:51:05 +08:00
XuFuandMick 924aeee59c [diffusion] feat: support batching for cosmos3 action generation (#36301)
Signed-off-by: FxxxxU <fu18801374388@163.com>
Signed-off-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-26 22:15:03 +08:00
Xinyuan Tong dfc40e0efe Add GLM-5.3-Flash cookbook (#36440) 2026-08-26 07:00:16 -07:00
Yuhao Yang 8eaffdf382 docs: point the Qwen3.8-Flash-Next cookbook at model support PR #36497 (#36499) 2026-08-26 12:53:55 +00:00
c7b5e76fa9 Add Qwen3.8-Flash-Next cookbook (#36496)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 20:36:59 +08:00
Mick c3c529c28e [diffusion] docs: distinguish MiniMax H3 checkpoint variants (#36412) 2026-08-26 15:16:18 +08:00
Yihao WangandClaude Opus 5 d7baad0116 [diffusion] Keep the Cosmos3 Super DiT resident on high-memory GPUs (#36375)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 21:10:39 -07:00
Mick 054f485d38 [diffusion] docs: add MiniMax H3 checkpoint format table (#36028) 2026-08-26 09:18:07 +08:00
Baizhou ZhangandRyan Stewart 41e7612dee [Model] Support Nemotron 3.5 Lightning speculative decoding (#36186)
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-25 16:43:58 -07:00
Junpan Wu 4b4bf3d2a5 [Deepseek-V4] Enable shared-experts fusion on the flashinfer_mxfp4 (trtllm-gen) MoE path (#35505)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
2026-08-25 14:55:25 -07:00
Xinyuan Tong 99c02d71b1 docs(cookbook): use auto parser resolution for Granite 4.2 (#36342) 2026-08-25 10:38:46 -07:00
Xinyuan Tong b760f7fb19 docs(cookbook): add IBM Granite 4.2 cookbook (#36286) 2026-08-25 22:54:55 +08:00