Commit Graph
113 Commits
Author SHA1 Message Date
Xinyuan Tong e27a7fac77 GLM-5.3-Flash cookbook: default Blackwell recipes to FP8 KV + TRT-LLM DSA (#36519) 2026-08-27 00:51:05 +08:00
XuFuandMick 924aeee59c [diffusion] feat: support batching for cosmos3 action generation (#36301)
Signed-off-by: FxxxxU <fu18801374388@163.com>
Signed-off-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-26 22:15:03 +08:00
Xinyuan Tong dfc40e0efe Add GLM-5.3-Flash cookbook (#36440) 2026-08-26 07:00:16 -07:00
Yuhao Yang 8eaffdf382 docs: point the Qwen3.8-Flash-Next cookbook at model support PR #36497 (#36499) 2026-08-26 12:53:55 +00:00
c7b5e76fa9 Add Qwen3.8-Flash-Next cookbook (#36496)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 20:36:59 +08:00
Mick c3c529c28e [diffusion] docs: distinguish MiniMax H3 checkpoint variants (#36412) 2026-08-26 15:16:18 +08:00
Yihao WangandClaude Opus 5 d7baad0116 [diffusion] Keep the Cosmos3 Super DiT resident on high-memory GPUs (#36375)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 21:10:39 -07:00
Mick 054f485d38 [diffusion] docs: add MiniMax H3 checkpoint format table (#36028) 2026-08-26 09:18:07 +08:00
Baizhou ZhangandRyan Stewart 41e7612dee [Model] Support Nemotron 3.5 Lightning speculative decoding (#36186)
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-25 16:43:58 -07:00
Junpan Wu 4b4bf3d2a5 [Deepseek-V4] Enable shared-experts fusion on the flashinfer_mxfp4 (trtllm-gen) MoE path (#35505)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
2026-08-25 14:55:25 -07:00
Xinyuan Tong 99c02d71b1 docs(cookbook): use auto parser resolution for Granite 4.2 (#36342) 2026-08-25 10:38:46 -07:00
Xinyuan Tong b760f7fb19 docs(cookbook): add IBM Granite 4.2 cookbook (#36286) 2026-08-25 22:54:55 +08:00
MickandClaude Fable 5 c3947eeada [diffusion] docs: desktop-safe 24 GB recipe and the DGX Spark tier (#36169)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 20:25:17 +08:00
a618d4c064 [AMD] Add Kimi-K2.7-Code-MXFP4 to cookbook (#36246)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-25 17:37:32 +08:00
Mick 191244b3f6 [diffusion] feat: support loading Comfy NVFP4-AWQ text encoders (#36046) 2026-08-25 15:43:30 +08:00
Артем Савкинandronnie_zheng 61b67316d8 [NPU] [Diffusion] Support MiniMax H3 on Ascend NPU's (#33569)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-25 09:44:14 +03:00
jacky.cheng 9eee990ce1 [AMD] cookbook: add HiCache host-DRAM KV tier for Qwen3.5 MXFP4 on MI355X (#36245) 2026-08-25 12:55:52 +08:00
Mick 67853c5804 [diffusion] feat: dispatch fp8 companions in mixed NVFP4 checkpoints (#36066) 2026-08-25 11:26:37 +08:00
Mick ddea7b9156 [diffusion] feat: support mixed Comfy NVFP4 and INT8 layers (#36061) 2026-08-25 09:19:47 +08:00
Xinyuan Tong 6e2f87d589 docs: mark Ling-3.0-flash DSPARK verified for all four quantizations on H200 (#36204) 2026-08-25 02:05:32 +08:00
Jimmy Shong 5030637c65 [docs] Split the Qwen3.8-27B NVFP4 cells by lm_head precision (#36020) 2026-08-25 01:50:33 +08:00
Mick 30f9ed09d1 [diffusion] feat: support loading minimax h3 gguf text encoders (#36055) 2026-08-24 23:00:49 +08:00
Mick 9b0007ed19 [diffusion] feat: support loading comfy nvfp4 minimax h3 checkpoints (#36044) 2026-08-24 22:57:02 +08:00
Mick 76d1401881 [diffusion] feat: support mixed w4a4 and int8 checkpoints (#36040) 2026-08-24 20:35:11 +08:00
Mick bfeae4e79a [diffusion] feat: support loading serialized convrot w4a4 checkpoints (#36039) 2026-08-24 19:13:35 +08:00
Mick adc09a1f63 [diffusion] feat: support loading self-describing quanto int8 encoders (#36052) 2026-08-24 16:37:09 +08:00
Xiaoyu Zhang 9866fe910b [diffusion] Speed up LingBot high-quality VAE decode (#36024) 2026-08-24 14:13:03 +08:00
Mick 8df3b9eff9 [diffusion] feat: support loading mixed w4a8 text encoders (#36037) 2026-08-24 13:31:27 +08:00
Xiaoyu Zhang 09592f5889 [diffusion] Keep LongLive2 components resident on large GPUs (#35993) 2026-08-24 12:06:52 +08:00
Mick 2d84de5e69 [diffusion] feat: support loading serialized comfy w4a8 checkpoints (#36036) 2026-08-24 10:32:21 +08:00
Mick 1c1c9d9b4e [diffusion] refactor: reuse srt quantization contracts and mxfp8 kernels (#36063) 2026-08-24 09:26:54 +08:00
Xiaoyu Zhang b2eb0fa51e [diffusion] Keep Cosmos3 Nano resident on high-memory GPUs (#36000) 2026-08-24 08:51:58 +08:00
Thomas Wang 95f5ecd3d2 [AMD] Update amd deepseek v4 cookbook 0822 (#35854) 2026-08-23 13:26:39 -07:00
amote-i 9b1b06b8e6 [NPU] [DOC] Add Ascend NPU (A3) recipe to the Kimi-K3 cookbook (#35508) 2026-08-23 21:21:27 +08:00
Mick dd15fb57b5 [diffusion] feat: automatically infer comfy fp8 activation scaling (#36060) 2026-08-23 18:43:33 +08:00
Mick a36c0746b9 [diffusion] feat: support serialized comfy convrot int8 dits (#35994) 2026-08-23 09:29:59 +08:00
MickandClaude Opus 5 7d22b7a875 [diffusion] docs: add tuning guide for h3 on consumer-level gpu (#35816)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 23:48:40 +08:00
Jianfei Wangandmiraclezqc af39ad9349 [Model] Complete dots.note.omni support with native encoders, video preprocessing, and MTP decoding (#33829)
Co-authored-by: miraclezqc <dysania@pku.edu.cn>
2026-08-22 14:19:14 +08:00
Jimmy ShongandClaude Fable 5 4cb5aebfe0 [docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spark grids on 1cf2b8c (#35825)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 13:03:18 +08:00
Baizhou Zhang 3b5909de0e [DeepSeek V4] Add W4A4 MegaMoE server flag (#35918) 2026-08-21 18:44:18 -07:00
zijiexiaandClaude Opus 5 fe8f9d7457 [Docs] Add --prerelease=allow to cookbook uv install commands (#35920)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:51:48 -07:00
Xinyuan Tong 05c584c44f docs: add DSPARK speculative decoding option to Ling-3.0-flash cookbook (#35861) 2026-08-22 02:50:55 +08:00
Jimmy Shong 3efa057449 [docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c (#35786) 2026-08-20 21:54:51 -07:00
MickandClaude Opus 5 e0cf75d9bd [doc] standardize diffusion cookbook model pages (#34247)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 10:25:40 +08:00
Jimmy Shong 14795dcb1a [docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev image tag (#35767) 2026-08-20 23:43:55 +00:00
Jimmy Shong 1a138e13b9 [docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753) 2026-08-20 23:34:44 +00:00
Jimmy Shong d9f6861359 [docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663) 2026-08-20 13:26:55 -07:00
Mick 82c6fc2db9 [diffusion] quant: support pruned safetensors checkpoints for minimax-h3 (#35418) 2026-08-20 19:34:14 +08:00
21c88f8625 [diffusion] quant: support gguf (#35370)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-20 15:46:34 +08:00
Jason Wiemels defb2a3100 feat(openai): Accept the input_audio content part in chat completions (#33606) 2026-08-19 13:37:50 -07:00