Commit Graph
55 Commits
Author SHA1 Message Date
MickandClaude Opus 5 7d22b7a875 [diffusion] docs: add tuning guide for h3 on consumer-level gpu (#35816)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 23:48:40 +08:00
Jianfei Wangandmiraclezqc af39ad9349 [Model] Complete dots.note.omni support with native encoders, video preprocessing, and MTP decoding (#33829)
Co-authored-by: miraclezqc <dysania@pku.edu.cn>
2026-08-22 14:19:14 +08:00
Jimmy ShongandClaude Fable 5 4cb5aebfe0 [docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spark grids on 1cf2b8c (#35825)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 13:03:18 +08:00
Baizhou Zhang 3b5909de0e [DeepSeek V4] Add W4A4 MegaMoE server flag (#35918) 2026-08-21 18:44:18 -07:00
Xinyuan Tong 05c584c44f docs: add DSPARK speculative decoding option to Ling-3.0-flash cookbook (#35861) 2026-08-22 02:50:55 +08:00
Jimmy Shong 3efa057449 [docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c (#35786) 2026-08-20 21:54:51 -07:00
MickandClaude Opus 5 e0cf75d9bd [doc] standardize diffusion cookbook model pages (#34247)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 10:25:40 +08:00
Jimmy Shong 1a138e13b9 [docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753) 2026-08-20 23:34:44 +00:00
Jimmy Shong d9f6861359 [docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663) 2026-08-20 13:26:55 -07:00
Xinyuan Tong 157d8ad27a Support Intern-S2-Mobius FP8 (#34908) 2026-08-19 10:58:01 -07:00
jacky.cheng 574274660f [AMD] cookbook: serve Qwen3.5 MXFP4 on MI355X with an fp8_e4m3 KV cache (#35445) 2026-08-19 18:49:32 +08:00
MickandClaude Opus 5 77fc5c128e [perf] overlap page preprocessing, pack the vit, enable prefill CUDA graph for paddle-ocr (#35318)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 08:20:55 +08:00
Thomas Wang 27596abdc0 [AMD] Update amd k3 cookbook for PR#34580 (#35263) 2026-08-17 23:19:36 -07:00
Baizhou Zhang 53621818e4 [Docs] Enable PD disaggregation for DSV4 low-latency recipes (#35224) 2026-08-17 20:07:27 -07:00
Jimmy ShongandClaude Opus 5 b956e916ae docs(cookbook): add Qwen3.8-27B DGX Spark configs (#35121)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:04:17 -07:00
Yuhao Yang 861eca8e25 docs: add NVFP4 quantization option to Kimi-K3 deploy panel (#35168) 2026-08-17 11:01:47 -07:00
Jimmy ShongandClaude Fable 5 e03c53fc13 docs(cookbook): Qwen3.8-27B deployment grid rework (#35065)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 00:56:08 -07:00
Jimmy ShongandClaude Fable 5 07a28ec5cf docs: fix Qwen3.8-27B mamba ratio calculator for speculative decoding (#35064)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 22:54:57 -07:00
zijiexiaandClaude Opus 5 f019f0b064 [Docs] Feature MiniMax-H3 in the popular-models banner (#35068)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 02:09:26 +00:00
5c0ace30c0 [diffusion] model: support ltx-2.5 (#34471)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-15 23:36:02 +08:00
Han-Yin ChangandClaude Fable 5 22dde1dd5b [Docs] Fill GLM-5.2 H200 FP8 speed cells (low-latency, balanced); fix MTP notation (#31554)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:56:55 -07:00
70e291b70f [Docs] Add GB300 cells and benchmarks for Qwen3.8-27B (#34863)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:42:50 +00:00
29c6be15a4 [Docs] Add Qwen3.8-27B cookbook page (#34860)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:23:39 +00:00
zijiexiaandClaude Opus 5 463981922c [Cookbook] Add DeepSeek-V4-Pro-0813 (Pro Official) serving recipes (#34809)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:53:41 -07:00
Yuhao YangandClaude 6ad3f2d8fd docs: link dots3.note checkpoints, add H100 cells (#34797)
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-14 02:42:27 +00:00
zijiexia abdef3c38e [Qwen] Update Docker image tag for MI300X to v0.5.17-rocm700-mi30x-20… (#34770) 2026-08-13 15:59:42 -07:00
c1142677a8 [do not merge] add new cookbooks (#34658)
Co-authored-by: Jianfei Wang <jianfei.wangg@outlook.com>
Co-authored-by: Jianfei Wang <905787410@qq.com>
2026-08-13 13:35:15 +00:00
Jimmy Shong 198b7e9240 [Docs] Use Meta's canonical Muse Glimmer GGUF filename (#34626) 2026-08-12 16:07:55 -07:00
YAMY e6250c7c70 docs: update Qwen3.8 disaggregated serving configs (#34601) 2026-08-12 11:02:35 -07:00
Xinyuan Tong d21eefc94f [Docs] Rename Qwen3.8-Max-DSpark to Qwen3.8-2.4T-A95B-DSpark (#34590) 2026-08-12 15:25:48 +00:00
Yichi Zhang f28bc5a6de docs(cookbook): add BF16 recipes to Nemotron 3.5 Lightning (#34573) 2026-08-12 08:15:34 -07:00
zijiexiaandClaude Opus 5 8e7c07fae7 [Docs] Add Qwen3.8 cookbook (#34587)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 08:07:46 -07:00
ajith-sirra-amdandgiovanniguastiamd b5d1453ed2 [AMD] GLM 5.2 MXFP4 SGLANG COOKBOOK (#34379)
Signed-off-by: Sirra <asirra@amd.com>
Co-authored-by: giovanniguastiamd <giovanni.guasti@amd.com>
2026-08-12 03:39:25 -07:00
Mohammad Miadh Angkad 00e57d74f0 Bump FlashInfer to 0.6.17 and remove Kimi K3 workarounds (#33997) 2026-08-12 02:17:26 -07:00
Faradawn Yang 857910bd35 docs(semianalysis): Update Qwen3.5 B200 NVFP4 MTP config (#34357) 2026-08-11 16:12:17 -07:00
Douglas YangandClaude Opus 5 59450c4f18 docs(cookbook): Kimi-K3 — drop --enable-symm-mem from the GB cells (#34444)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 19:23:41 +00:00
Xinyuan Tong d5d41d07ed [Docs] Add Ling-3.0-tiny INT4 recipes (#34395) 2026-08-12 03:16:03 +08:00
Faradawn YangandRyan Stewart 3add7e19ff Add NVIDIA Nemotron 3.5 Lightning cookbook (#33481)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
Signed-off-by: Ryan Stewart <rystewart@nvidia.com>
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-11 06:00:57 -07:00
Xinyuan Tong 1c06c160f9 [Docs] Add Ling-3.0-flash INT4 and MXFP4 recipes (#34363) 2026-08-11 00:35:27 -07:00
Baizhou Zhang a92bbf2f24 docs: remove DSV4 low-latency chunked prefill size (#34333) 2026-08-10 17:23:43 -07:00
Brayden ZhongandBrayden Zhong b86c90215f [Docs] Muse Glimmer cookbook: drop --speculative-dflash-block-size 5 (#34323)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-10 15:46:34 -07:00
e54c153ba6 Add Intern-S2-Mobius cookbook (#33820)
Co-authored-by: Justin Tong <justintong0323@outlook.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-10 13:13:35 -07:00
Xinyuan Tong 77c90e7e54 Cookbook: add Ling-3.0-tiny (#34283) 2026-08-10 21:25:17 +08:00
a6c34df044 Muse Glimmer Cookbook (#34271)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: Jimmy Shong <jimmysh341@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-10 10:21:13 +00:00
Ke Bao 430f38ea25 Update dspark draft path in Inkling small cookbook (#34250) 2026-08-10 17:12:51 +08:00
Mick cf2d4fd679 docs: clarify K3 VLM feature transport (#34099) 2026-08-08 19:23:00 +08:00
MickandClaude Fable 5 a25c330eb1 [diffusion] feat: cross-node sequence parallelism (Ulysses x Ring) (#33327)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 10:56:28 +08:00
Douglas YangandClaude Opus 5 86f373daff docs(cookbook): DeepSeek-V4-Flash-0731 — drop chunked-prefill/autotune flags on B300 low-latency (#34044)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 23:12:56 +00:00
Faradawn Yang 115cd7bde1 docs: update checkpoint to Qwen3.5 NVFP4 V2 for InfX (#32945) 2026-08-07 15:26:07 -07:00
Xinyuan Tong 0da25ee6f7 Docs: Ling-3.0-flash cookbook — serve native 256K, drop YaRN override (#33882) 2026-08-07 21:46:51 +00:00