Commit Graph
1610 Commits
Author SHA1 Message Date
triple-muandmickqian bf71035d39 [diffusion] MiniMax-H3: tiered AdaLN plan cache (pinned-host tier + per-plan LRU) (#37266)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
2026-09-03 22:09:21 +08:00
2bb25dc18b [Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-03 20:08:41 +08:00
amote-i 354ed6d66b [NPU] [DOC] Refresh supported models and features on NPU (#37799) 2026-09-03 19:53:39 +08:00
kkandwunhuang dd091f43cd [AMD] Update kimi-k3 amd cookbook 0903 (#37781)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-03 18:44:39 +08:00
Yash Akhauri 02d9b3060a [Docs] Update K2 Horizon MoE model names (#37723) 2026-09-02 23:32:44 -07:00
Yash Akhauri 98ef7d8ae6 docs: add K2 Horizon cookbook recipes and H200 results (#37655) 2026-09-03 11:52:05 +08:00
James Liu 4229088a48 feat(kernels): generalize persistent CuTe JIT cache (#33911) 2026-09-02 19:59:04 -07:00
Alex NailsandAlison Shao 28262c20df [CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-09-02 19:46:08 -07:00
Oguz Ulgen f15748d965 [bench] Support real-traffic replay with early-stop-aware steady-state metrics in bench_one_batch_server (#37469) 2026-09-02 17:20:25 -07:00
Xinyuan Tong 3421d4375b [Docs] GLM-5.3-Flash cookbook: drop stale EP caveat, add B300/H100/B200 FP8 speed data (#37576) 2026-09-02 17:16:23 -07:00
Faradawn Yang 9c70d22721 Update GLM-5.2 NVFP4 B200/B300 for AgentX HiCache (#35368) 2026-09-02 14:39:09 -07:00
Mick f6aed6ec53 [diffusion] doc: rewrite stale diffusion compatibility matrix (#36987) 2026-09-02 23:42:54 +08:00
Kevin MiandClaude Fable 5 f586654518 [diffusion] feat: support FastH3 (4-step VSA-distilled MiniMax-H3) with a VSA-H3 attention backend (#37480)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-02 21:39:54 +08:00
Liangsheng Yin ebfd8c60e5 [CI] Install sgl-eval from PyPI through the test extra (#37504) 2026-09-02 01:45:25 -07:00
4b329482e8 [diffusion] feat: support cube sparse attention for minimax h3 (#34893)
Co-authored-by: zhenaozhenfu <zhenaozhenfu@minimaxi.com>
Co-authored-by: Reynor <reynor@minimaxi.com>
2026-09-02 15:21:33 +08:00
Mick 9175590aa0 [diffusion] refactor: admit explicit attention backends by capability (#37441) 2026-09-02 15:17:57 +08:00
Xiaoyu Zhang 1aa8299d1d [Diffusion] Add cumulative extra-high quality tier (#37422) 2026-09-02 10:26:13 +08:00
Mick dde0ecdb90 [diffusion] feat: support spargeattention (#37437) 2026-09-02 09:42:33 +08:00
zijiexia 6d34a4d3ce [Cookbook] Verify DeepSeek-V4 Flash Vision on GB300 (#37492) 2026-09-01 17:11:41 -07:00
Jimmy ShongandClaude Fable 5.1 ed82bea146 [Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-01 15:42:39 -07:00
zijiexiaandClaude Fable 5 0f18d389b4 [Cookbook] Verify DeepSeek-V4 Flash Vision balanced and high-throughput on B200 (#37468)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 13:09:14 -07:00
Xinyuan Tong 442c7c1e29 [Docs] GLM-5.3-Flash cookbook: add NVFP4 FP8+TRT-LLM benchmark rows (follow-up to #37109) (#37412) 2026-09-01 12:56:22 -07:00
Ankur Singh 3315356cc0 docs(cookbook): enable FlashInfer GDN for Qwen3.5 B200 (#37360) 2026-09-01 11:48:41 -07:00
YC Yen-Ching Tseng b425897366 [AMD] Gate the aiter memory-reserve exemption behind an env var (#37242) 2026-09-01 03:46:15 -07:00
zijiexiaandClaude Opus 5 dc1ae02684 [Cookbook] Add the DFlash2 speculative option to GLM-5.3 (#37392)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 09:08:28 +00:00
zijiexia 6c72b49a57 Revert "[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)" (#37380) 2026-09-01 01:25:13 -07:00
zijiexiaandClaude Fable 5 379e33d87e [Cookbook] Add NVFP4 options for DeepSeek-V4 Flash Official (0731) and Pro Official (0813) (#37351)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 08:03:39 +00:00
Xinyuan Tong 60548501bb [Docs] Add NVFP4 section to GLM-5.3-Flash cookbook (#37109) 2026-09-01 14:14:30 +08:00
Yuan Luoandluoyuan.luo 5b04408784 [MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE (#34967)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-08-31 20:04:41 -07:00
71cee04ebe [Diffusion] Optimize Qwen-Image TP collectives and attention (#36680)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:28:37 +08:00
Ziang Li 9a85473a89 [FlashInfer v0.6.18] add FlashInfer CuTe DSL NVFP4 W4A16 mode (#35120) 2026-08-31 18:47:30 -07:00
zijiexia 455232de6e [Cookbook] Enable DSpark on the DeepSeek-V4 Flash Vision low-latency recipes (#37301) 2026-08-31 16:43:54 -07:00
zijiexia 88cf5c9541 [Cookbook] Add DeepSeek-V4-Flash-Vision-Exp to the DeepSeek-V4 page (#37293) 2026-08-31 14:52:28 -07:00
Mick 62c470697e [diffusion] chore: enforce component attention backend application (#36907) 2026-08-31 14:00:02 +08:00
Mick 881cbfe54c [diffusion] feat: add exact component precision overrides (#36991) 2026-08-31 11:12:07 +08:00
Mohammad Miadh Angkad 4f761e8649 [Deps] Bump FlashInfer to 0.6.18 (#36954) 2026-08-30 19:02:39 -07:00
Mick fe694986a2 [diffusion] chore: make malformed component execution options fail-fast (#37049) 2026-08-30 21:06:19 +08:00
WenhaoZhang e9a7157615 [diffusion] feat: allow cache-dit with dit layerwise offload (#35858) 2026-08-30 20:55:41 +08:00
Mick aa483ab782 [diffusion] feat: support streaming native vae weights directly to gpu (#37004) 2026-08-30 20:48:09 +08:00
Thomas Wang 7399c2b558 [AMD] Update v4 amd cookbook 0830 (#37092) 2026-08-29 23:55:28 -07:00
Liangsheng Yin 9a489f8d2f [Test] Move gpqa and aime25 onto sgl-eval, drop unused eval paths (#36979) 2026-08-29 17:36:13 -07:00
Shuwen WangandClaude Opus 5 000c636342 docs: state that HiCache L2 is instance-private and only L3 is shared (#37050)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 23:47:50 +08:00
Mick b24bd44556 [diffusion] feat: avoid direct GPU parameter copies (#36832) 2026-08-29 22:43:25 +08:00
Xiaoyu Zhang 0e1146d04f [diffusion] optimization: optimize Pi0.5 inference and bounded graph serving (#34599) 2026-08-29 14:47:32 +08:00
Mick fa474b0441 [diffusion] fix: fix image encoder parallel folding proposal (#36863) 2026-08-29 14:34:36 +08:00
Liangsheng Yin a25df83fe3 [Cookbook] Run accuracy benchmarks through sgl-eval (#36977) 2026-08-28 23:29:38 -07:00
Mick d1ce017665 [diffusion] feat: delegate recognized quantized components to transformers (#36902) 2026-08-29 14:26:50 +08:00
Mohammad Miadh Angkad 8a4c517a60 [Docs] Restore the AIME25 label so GLM-5.3 FP8 and BF16 scores render again (#36950) 2026-08-29 11:38:26 +08:00
amote-i 505228823f [NPU] [DOC] udpate supported features on NPU (#36940) 2026-08-29 10:52:01 +08:00
amote-i 51c18d9aa8 [NPU] [DOC] update npu best practice (#36476) 2026-08-29 09:38:52 +08:00