Mick
|
dde0ecdb90
|
[diffusion] feat: support spargeattention (#37437)
|
2026-09-02 09:42:33 +08:00 |
|
zijiexia
|
6d34a4d3ce
|
[Cookbook] Verify DeepSeek-V4 Flash Vision on GB300 (#37492)
|
2026-09-01 17:11:41 -07:00 |
|
 Jimmy ShongandClaude Fable 5.1
|
ed82bea146
|
[Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-01 15:42:39 -07:00 |
|
 zijiexiaandClaude Fable 5
|
0f18d389b4
|
[Cookbook] Verify DeepSeek-V4 Flash Vision balanced and high-throughput on B200 (#37468)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-01 13:09:14 -07:00 |
|
Xinyuan Tong
|
442c7c1e29
|
[Docs] GLM-5.3-Flash cookbook: add NVFP4 FP8+TRT-LLM benchmark rows (follow-up to #37109) (#37412)
|
2026-09-01 12:56:22 -07:00 |
|
Ankur Singh
|
3315356cc0
|
docs(cookbook): enable FlashInfer GDN for Qwen3.5 B200 (#37360)
|
2026-09-01 11:48:41 -07:00 |
|
YC Yen-Ching Tseng
|
b425897366
|
[AMD] Gate the aiter memory-reserve exemption behind an env var (#37242)
|
2026-09-01 03:46:15 -07:00 |
|
 zijiexiaandClaude Opus 5
|
dc1ae02684
|
[Cookbook] Add the DFlash2 speculative option to GLM-5.3 (#37392)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-01 09:08:28 +00:00 |
|
zijiexia
|
6c72b49a57
|
Revert "[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)" (#37380)
|
2026-09-01 01:25:13 -07:00 |
|
 zijiexiaandClaude Fable 5
|
379e33d87e
|
[Cookbook] Add NVFP4 options for DeepSeek-V4 Flash Official (0731) and Pro Official (0813) (#37351)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-01 08:03:39 +00:00 |
|
Xinyuan Tong
|
60548501bb
|
[Docs] Add NVFP4 section to GLM-5.3-Flash cookbook (#37109)
|
2026-09-01 14:14:30 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
5b04408784
|
[MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE (#34967)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-08-31 20:04:41 -07:00 |
|
 
|
71cee04ebe
|
[Diffusion] Optimize Qwen-Image TP collectives and attention (#36680)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 10:28:37 +08:00 |
|
Ziang Li
|
9a85473a89
|
[FlashInfer v0.6.18] add FlashInfer CuTe DSL NVFP4 W4A16 mode (#35120)
|
2026-08-31 18:47:30 -07:00 |
|
zijiexia
|
455232de6e
|
[Cookbook] Enable DSpark on the DeepSeek-V4 Flash Vision low-latency recipes (#37301)
|
2026-08-31 16:43:54 -07:00 |
|
zijiexia
|
88cf5c9541
|
[Cookbook] Add DeepSeek-V4-Flash-Vision-Exp to the DeepSeek-V4 page (#37293)
|
2026-08-31 14:52:28 -07:00 |
|
Mick
|
62c470697e
|
[diffusion] chore: enforce component attention backend application (#36907)
|
2026-08-31 14:00:02 +08:00 |
|
Mick
|
881cbfe54c
|
[diffusion] feat: add exact component precision overrides (#36991)
|
2026-08-31 11:12:07 +08:00 |
|
Mohammad Miadh Angkad
|
4f761e8649
|
[Deps] Bump FlashInfer to 0.6.18 (#36954)
|
2026-08-30 19:02:39 -07:00 |
|
Mick
|
fe694986a2
|
[diffusion] chore: make malformed component execution options fail-fast (#37049)
|
2026-08-30 21:06:19 +08:00 |
|
WenhaoZhang
|
e9a7157615
|
[diffusion] feat: allow cache-dit with dit layerwise offload (#35858)
|
2026-08-30 20:55:41 +08:00 |
|
Mick
|
aa483ab782
|
[diffusion] feat: support streaming native vae weights directly to gpu (#37004)
|
2026-08-30 20:48:09 +08:00 |
|
Thomas Wang
|
7399c2b558
|
[AMD] Update v4 amd cookbook 0830 (#37092)
|
2026-08-29 23:55:28 -07:00 |
|
Liangsheng Yin
|
9a489f8d2f
|
[Test] Move gpqa and aime25 onto sgl-eval, drop unused eval paths (#36979)
|
2026-08-29 17:36:13 -07:00 |
|
 Shuwen WangandClaude Opus 5
|
000c636342
|
docs: state that HiCache L2 is instance-private and only L3 is shared (#37050)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-29 23:47:50 +08:00 |
|
Mick
|
b24bd44556
|
[diffusion] feat: avoid direct GPU parameter copies (#36832)
|
2026-08-29 22:43:25 +08:00 |
|
Xiaoyu Zhang
|
0e1146d04f
|
[diffusion] optimization: optimize Pi0.5 inference and bounded graph serving (#34599)
|
2026-08-29 14:47:32 +08:00 |
|
Mick
|
fa474b0441
|
[diffusion] fix: fix image encoder parallel folding proposal (#36863)
|
2026-08-29 14:34:36 +08:00 |
|
Liangsheng Yin
|
a25df83fe3
|
[Cookbook] Run accuracy benchmarks through sgl-eval (#36977)
|
2026-08-28 23:29:38 -07:00 |
|
Mick
|
d1ce017665
|
[diffusion] feat: delegate recognized quantized components to transformers (#36902)
|
2026-08-29 14:26:50 +08:00 |
|
Mohammad Miadh Angkad
|
8a4c517a60
|
[Docs] Restore the AIME25 label so GLM-5.3 FP8 and BF16 scores render again (#36950)
|
2026-08-29 11:38:26 +08:00 |
|
amote-i
|
505228823f
|
[NPU] [DOC] udpate supported features on NPU (#36940)
|
2026-08-29 10:52:01 +08:00 |
|
amote-i
|
51c18d9aa8
|
[NPU] [DOC] update npu best practice (#36476)
|
2026-08-29 09:38:52 +08:00 |
|
Thomas Wang
|
89816a21a1
|
[AMD] Update v4 amd cookbook 0828 (#36828)
|
2026-08-28 17:20:22 -07:00 |
|
Xiaoyu Zhang
|
db6f0a9d53
|
Refactor JIT kernel and expert-pack directory layout (#36704)
|
2026-08-29 07:41:25 +08:00 |
|
Xiaoyu Zhang
|
50bc1a3767
|
[diffusion] Keep Cosmos3 Nano resident on 96 GB GPUs (#36641)
|
2026-08-29 07:40:46 +08:00 |
|
  
|
395c2258c3
|
[Docs] Add GLM-5.3 cookbook (#36827)
Co-authored-by: JustinTong0323 <xinyuantong.cs@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Mohammad Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-08-28 14:56:06 +00:00 |
|
 Артем СавкинandXiaoyu Zhang
|
ecbadf0b4b
|
[NPU] [Diffusion] support distributed inference pipeline for GLM-Image (#31320)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-28 15:39:05 +03:00 |
|
Mick
|
803b4fb31c
|
[diffusion] refactor: scope model-specific API parameters (#35613)
|
2026-08-28 19:08:30 +08:00 |
|
Haoguang Cai
|
989e51ba9c
|
[Docs] Rename Tencent cookbook page titles to "Hy4 preview" / "Hy3 preview" (#36823)
|
2026-08-28 01:31:52 -07:00 |
|
 zijiexiaandClaude Fable 5
|
2960d69622
|
[Cookbook] Hy4-Preview follow-ups: runtime-accurate recipes + released-model info (#36808)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-28 00:19:07 -07:00 |
|
 zijiexiaandClaude Fable 5
|
1948b61ad4
|
[Cookbook] Add the Hy4-Preview model page (Tencent) (#36804)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-27 23:26:13 -07:00 |
|
 zijiexiaandClaude Fable 5
|
43b5a57dbb
|
[Docs] Feature GLM-5.3-Flash in the popular-models banner (#36784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-27 21:56:10 -07:00 |
|
 zijiexiaandClaude Opus 5
|
6ccfeb59bc
|
cookbook: add a Speculative card to the GLM-5.3-Flash playground (#36740)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-27 16:03:59 -07:00 |
|
Xinyuan Tong
|
d1f14431fd
|
GLM-5.3-Flash cookbook: HiCache for LL, fusion-flag drop, EAGLE, default-cell numbers, DCP4 overlay (#36544)
|
2026-08-28 03:18:25 +08:00 |
|
 zijiexiaandClaude Opus 5
|
46a544e0a0
|
[Docs] GLM-5.3-Flash: point at compute-mamba-ratio for the KDA/KV pool split (#36719)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-27 12:16:35 -07:00 |
|
Xinyuan Tong
|
11de5e2281
|
docs(cookbook): add GB10 (DGX Spark) MXFP4 cells for Ling-3.0-flash (#36364)
|
2026-08-27 12:15:17 -07:00 |
|
 
|
97ba99067d
|
Publish per-scheduler load on a dedicated socket for load-aware routers (#34608)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-08-27 19:33:18 +08:00 |
|
Julian Huang
|
536f570e66
|
docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state (#36611)
|
2026-08-27 02:49:03 -07:00 |
|
zijiexia
|
636a6f7dba
|
cookbook: fix GLM-5.3-Flash speculative flag, size Hopper memory, record GSM8K (#36660)
|
2026-08-27 02:19:47 -07:00 |
|