Commit Graph
1577 Commits
Author SHA1 Message Date
Mick 62c470697e [diffusion] chore: enforce component attention backend application (#36907) 2026-08-31 14:00:02 +08:00
Mick 881cbfe54c [diffusion] feat: add exact component precision overrides (#36991) 2026-08-31 11:12:07 +08:00
Mohammad Miadh Angkad 4f761e8649 [Deps] Bump FlashInfer to 0.6.18 (#36954) 2026-08-30 19:02:39 -07:00
Mick fe694986a2 [diffusion] chore: make malformed component execution options fail-fast (#37049) 2026-08-30 21:06:19 +08:00
WenhaoZhang e9a7157615 [diffusion] feat: allow cache-dit with dit layerwise offload (#35858) 2026-08-30 20:55:41 +08:00
Mick aa483ab782 [diffusion] feat: support streaming native vae weights directly to gpu (#37004) 2026-08-30 20:48:09 +08:00
Thomas Wang 7399c2b558 [AMD] Update v4 amd cookbook 0830 (#37092) 2026-08-29 23:55:28 -07:00
Liangsheng Yin 9a489f8d2f [Test] Move gpqa and aime25 onto sgl-eval, drop unused eval paths (#36979) 2026-08-29 17:36:13 -07:00
Shuwen WangandClaude Opus 5 000c636342 docs: state that HiCache L2 is instance-private and only L3 is shared (#37050)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 23:47:50 +08:00
Mick b24bd44556 [diffusion] feat: avoid direct GPU parameter copies (#36832) 2026-08-29 22:43:25 +08:00
Xiaoyu Zhang 0e1146d04f [diffusion] optimization: optimize Pi0.5 inference and bounded graph serving (#34599) 2026-08-29 14:47:32 +08:00
Mick fa474b0441 [diffusion] fix: fix image encoder parallel folding proposal (#36863) 2026-08-29 14:34:36 +08:00
Liangsheng Yin a25df83fe3 [Cookbook] Run accuracy benchmarks through sgl-eval (#36977) 2026-08-28 23:29:38 -07:00
Mick d1ce017665 [diffusion] feat: delegate recognized quantized components to transformers (#36902) 2026-08-29 14:26:50 +08:00
Mohammad Miadh Angkad 8a4c517a60 [Docs] Restore the AIME25 label so GLM-5.3 FP8 and BF16 scores render again (#36950) 2026-08-29 11:38:26 +08:00
amote-i 505228823f [NPU] [DOC] udpate supported features on NPU (#36940) 2026-08-29 10:52:01 +08:00
amote-i 51c18d9aa8 [NPU] [DOC] update npu best practice (#36476) 2026-08-29 09:38:52 +08:00
Thomas Wang 89816a21a1 [AMD] Update v4 amd cookbook 0828 (#36828) 2026-08-28 17:20:22 -07:00
Xiaoyu Zhang db6f0a9d53 Refactor JIT kernel and expert-pack directory layout (#36704) 2026-08-29 07:41:25 +08:00
Xiaoyu Zhang 50bc1a3767 [diffusion] Keep Cosmos3 Nano resident on 96 GB GPUs (#36641) 2026-08-29 07:40:46 +08:00
395c2258c3 [Docs] Add GLM-5.3 cookbook (#36827)
Co-authored-by: JustinTong0323 <xinyuantong.cs@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Mohammad Angkad <176301910+mmangkad@users.noreply.github.com>
2026-08-28 14:56:06 +00:00
Артем СавкинandXiaoyu Zhang ecbadf0b4b [NPU] [Diffusion] support distributed inference pipeline for GLM-Image (#31320)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-08-28 15:39:05 +03:00
Mick 803b4fb31c [diffusion] refactor: scope model-specific API parameters (#35613) 2026-08-28 19:08:30 +08:00
Haoguang Cai 989e51ba9c [Docs] Rename Tencent cookbook page titles to "Hy4 preview" / "Hy3 preview" (#36823) 2026-08-28 01:31:52 -07:00
zijiexiaandClaude Fable 5 2960d69622 [Cookbook] Hy4-Preview follow-ups: runtime-accurate recipes + released-model info (#36808)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 00:19:07 -07:00
zijiexiaandClaude Fable 5 1948b61ad4 [Cookbook] Add the Hy4-Preview model page (Tencent) (#36804)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 23:26:13 -07:00
zijiexiaandClaude Fable 5 43b5a57dbb [Docs] Feature GLM-5.3-Flash in the popular-models banner (#36784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 21:56:10 -07:00
zijiexiaandClaude Opus 5 6ccfeb59bc cookbook: add a Speculative card to the GLM-5.3-Flash playground (#36740)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 16:03:59 -07:00
Xinyuan Tong d1f14431fd GLM-5.3-Flash cookbook: HiCache for LL, fusion-flag drop, EAGLE, default-cell numbers, DCP4 overlay (#36544) 2026-08-28 03:18:25 +08:00
zijiexiaandClaude Opus 5 46a544e0a0 [Docs] GLM-5.3-Flash: point at compute-mamba-ratio for the KDA/KV pool split (#36719)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 12:16:35 -07:00
Xinyuan Tong 11de5e2281 docs(cookbook): add GB10 (DGX Spark) MXFP4 cells for Ling-3.0-flash (#36364) 2026-08-27 12:15:17 -07:00
97ba99067d Publish per-scheduler load on a dedicated socket for load-aware routers (#34608)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-08-27 19:33:18 +08:00
Julian Huang 536f570e66 docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state (#36611) 2026-08-27 02:49:03 -07:00
zijiexia 636a6f7dba cookbook: fix GLM-5.3-Flash speculative flag, size Hopper memory, record GSM8K (#36660) 2026-08-27 02:19:47 -07:00
andyluo7 0f7b5b8b2a [AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608) 2026-08-27 05:19:57 +00:00
sglang-botandsglang-bot e6c1a3cd4c docs: sync LMSYS SGLang blog cards (#35416)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-27 03:56:54 +00:00
Xinyuan Tong e27a7fac77 GLM-5.3-Flash cookbook: default Blackwell recipes to FP8 KV + TRT-LLM DSA (#36519) 2026-08-27 00:51:05 +08:00
Xinyuan Tong f8cc1f9525 GLM-5.3-Flash cookbook: FP8 KV + TRT-LLM DSA benchmark card (#36513) 2026-08-26 07:39:27 -07:00
XuFuandMick 924aeee59c [diffusion] feat: support batching for cosmos3 action generation (#36301)
Signed-off-by: FxxxxU <fu18801374388@163.com>
Signed-off-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-26 22:15:03 +08:00
Xinyuan Tong dfc40e0efe Add GLM-5.3-Flash cookbook (#36440) 2026-08-26 07:00:16 -07:00
Yuhao Yang 8eaffdf382 docs: point the Qwen3.8-Flash-Next cookbook at model support PR #36497 (#36499) 2026-08-26 12:53:55 +00:00
c7b5e76fa9 Add Qwen3.8-Flash-Next cookbook (#36496)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 20:36:59 +08:00
Mick c3c529c28e [diffusion] docs: distinguish MiniMax H3 checkpoint variants (#36412) 2026-08-26 15:16:18 +08:00
amote-i abe3aeb142 [NPU] [DOC] Update CANN version in NPU installation docs (#36325) 2026-08-26 14:51:43 +08:00
Ziang Li 3c9febc68b [Spec][DSA] Add --speculative-dsa-topk-backend (#36313) 2026-08-25 23:35:03 -07:00
Yihao WangandClaude Opus 5 d7baad0116 [diffusion] Keep the Cosmos3 Super DiT resident on high-memory GPUs (#36375)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 21:10:39 -07:00
Alison Shao 3ec22948c1 [CI] /rerun-failed-ci: rerun cancelled runs and target the newest run per workflow (#34057) 2026-08-25 20:15:54 -07:00
34de1fb47f fix(test): stabilize nightly precision regression (#34668)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-08-25 20:04:52 -07:00
Mick 054f485d38 [diffusion] docs: add MiniMax H3 checkpoint format table (#36028) 2026-08-26 09:18:07 +08:00
Baizhou ZhangandRyan Stewart 41e7612dee [Model] Support Nemotron 3.5 Lightning speculative decoding (#36186)
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-25 16:43:58 -07:00