Liangsheng Yin
|
ebfd8c60e5
|
[CI] Install sgl-eval from PyPI through the test extra (#37504)
|
2026-09-02 01:45:25 -07:00 |
|
 Jimmy ShongandClaude Fable 5.1
|
ed82bea146
|
[Cookbook] DeepSeek-V4: add DGX Spark (2x GB10) Flash Official FP4 recipe (#37479)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-01 15:42:39 -07:00 |
|
 zijiexiaandClaude Opus 5
|
dc1ae02684
|
[Cookbook] Add the DFlash2 speculative option to GLM-5.3 (#37392)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-01 09:08:28 +00:00 |
|
zijiexia
|
6c72b49a57
|
Revert "[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)" (#37380)
|
2026-09-01 01:25:13 -07:00 |
|
 zijiexiaandClaude Fable 5
|
379e33d87e
|
[Cookbook] Add NVFP4 options for DeepSeek-V4 Flash Official (0731) and Pro Official (0813) (#37351)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-01 08:03:39 +00:00 |
|
Xinyuan Tong
|
60548501bb
|
[Docs] Add NVFP4 section to GLM-5.3-Flash cookbook (#37109)
|
2026-09-01 14:14:30 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
5b04408784
|
[MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE (#34967)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-08-31 20:04:41 -07:00 |
|
zijiexia
|
455232de6e
|
[Cookbook] Enable DSpark on the DeepSeek-V4 Flash Vision low-latency recipes (#37301)
|
2026-08-31 16:43:54 -07:00 |
|
zijiexia
|
88cf5c9541
|
[Cookbook] Add DeepSeek-V4-Flash-Vision-Exp to the DeepSeek-V4 page (#37293)
|
2026-08-31 14:52:28 -07:00 |
|
Mohammad Miadh Angkad
|
4f761e8649
|
[Deps] Bump FlashInfer to 0.6.18 (#36954)
|
2026-08-30 19:02:39 -07:00 |
|
Thomas Wang
|
7399c2b558
|
[AMD] Update v4 amd cookbook 0830 (#37092)
|
2026-08-29 23:55:28 -07:00 |
|
Liangsheng Yin
|
a25df83fe3
|
[Cookbook] Run accuracy benchmarks through sgl-eval (#36977)
|
2026-08-28 23:29:38 -07:00 |
|
Thomas Wang
|
89816a21a1
|
[AMD] Update v4 amd cookbook 0828 (#36828)
|
2026-08-28 17:20:22 -07:00 |
|
  
|
395c2258c3
|
[Docs] Add GLM-5.3 cookbook (#36827)
Co-authored-by: JustinTong0323 <xinyuantong.cs@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Mohammad Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-08-28 14:56:06 +00:00 |
|
Haoguang Cai
|
989e51ba9c
|
[Docs] Rename Tencent cookbook page titles to "Hy4 preview" / "Hy3 preview" (#36823)
|
2026-08-28 01:31:52 -07:00 |
|
 zijiexiaandClaude Fable 5
|
2960d69622
|
[Cookbook] Hy4-Preview follow-ups: runtime-accurate recipes + released-model info (#36808)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-28 00:19:07 -07:00 |
|
 zijiexiaandClaude Fable 5
|
1948b61ad4
|
[Cookbook] Add the Hy4-Preview model page (Tencent) (#36804)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-27 23:26:13 -07:00 |
|
 zijiexiaandClaude Fable 5
|
43b5a57dbb
|
[Docs] Feature GLM-5.3-Flash in the popular-models banner (#36784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-27 21:56:10 -07:00 |
|
 zijiexiaandClaude Opus 5
|
6ccfeb59bc
|
cookbook: add a Speculative card to the GLM-5.3-Flash playground (#36740)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-27 16:03:59 -07:00 |
|
Xinyuan Tong
|
d1f14431fd
|
GLM-5.3-Flash cookbook: HiCache for LL, fusion-flag drop, EAGLE, default-cell numbers, DCP4 overlay (#36544)
|
2026-08-28 03:18:25 +08:00 |
|
 zijiexiaandClaude Opus 5
|
46a544e0a0
|
[Docs] GLM-5.3-Flash: point at compute-mamba-ratio for the KDA/KV pool split (#36719)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-27 12:16:35 -07:00 |
|
zijiexia
|
636a6f7dba
|
cookbook: fix GLM-5.3-Flash speculative flag, size Hopper memory, record GSM8K (#36660)
|
2026-08-27 02:19:47 -07:00 |
|
andyluo7
|
0f7b5b8b2a
|
[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)
|
2026-08-27 05:19:57 +00:00 |
|
Xinyuan Tong
|
e27a7fac77
|
GLM-5.3-Flash cookbook: default Blackwell recipes to FP8 KV + TRT-LLM DSA (#36519)
|
2026-08-27 00:51:05 +08:00 |
|
Xinyuan Tong
|
dfc40e0efe
|
Add GLM-5.3-Flash cookbook (#36440)
|
2026-08-26 07:00:16 -07:00 |
|
Yuhao Yang
|
8eaffdf382
|
docs: point the Qwen3.8-Flash-Next cookbook at model support PR #36497 (#36499)
|
2026-08-26 12:53:55 +00:00 |
|
 
|
c7b5e76fa9
|
Add Qwen3.8-Flash-Next cookbook (#36496)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-26 20:36:59 +08:00 |
|
 Baizhou ZhangandRyan Stewart
|
41e7612dee
|
[Model] Support Nemotron 3.5 Lightning speculative decoding (#36186)
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
|
2026-08-25 16:43:58 -07:00 |
|
Junpan Wu
|
4b4bf3d2a5
|
[Deepseek-V4] Enable shared-experts fusion on the flashinfer_mxfp4 (trtllm-gen) MoE path (#35505)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
|
2026-08-25 14:55:25 -07:00 |
|
Xinyuan Tong
|
99c02d71b1
|
docs(cookbook): use auto parser resolution for Granite 4.2 (#36342)
|
2026-08-25 10:38:46 -07:00 |
|
Xinyuan Tong
|
b760f7fb19
|
docs(cookbook): add IBM Granite 4.2 cookbook (#36286)
|
2026-08-25 22:54:55 +08:00 |
|
 
|
a618d4c064
|
[AMD] Add Kimi-K2.7-Code-MXFP4 to cookbook (#36246)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-08-25 17:37:32 +08:00 |
|
jacky.cheng
|
9eee990ce1
|
[AMD] cookbook: add HiCache host-DRAM KV tier for Qwen3.5 MXFP4 on MI355X (#36245)
|
2026-08-25 12:55:52 +08:00 |
|
Xinyuan Tong
|
6e2f87d589
|
docs: mark Ling-3.0-flash DSPARK verified for all four quantizations on H200 (#36204)
|
2026-08-25 02:05:32 +08:00 |
|
Jimmy Shong
|
5030637c65
|
[docs] Split the Qwen3.8-27B NVFP4 cells by lm_head precision (#36020)
|
2026-08-25 01:50:33 +08:00 |
|
Thomas Wang
|
95f5ecd3d2
|
[AMD] Update amd deepseek v4 cookbook 0822 (#35854)
|
2026-08-23 13:26:39 -07:00 |
|
amote-i
|
9b1b06b8e6
|
[NPU] [DOC] Add Ascend NPU (A3) recipe to the Kimi-K3 cookbook (#35508)
|
2026-08-23 21:21:27 +08:00 |
|
 Jianfei Wangandmiraclezqc
|
af39ad9349
|
[Model] Complete dots.note.omni support with native encoders, video preprocessing, and MTP decoding (#33829)
Co-authored-by: miraclezqc <dysania@pku.edu.cn>
|
2026-08-22 14:19:14 +08:00 |
|
 Jimmy ShongandClaude Fable 5
|
4cb5aebfe0
|
[docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spark grids on 1cf2b8c (#35825)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-22 13:03:18 +08:00 |
|
Baizhou Zhang
|
3b5909de0e
|
[DeepSeek V4] Add W4A4 MegaMoE server flag (#35918)
|
2026-08-21 18:44:18 -07:00 |
|
 zijiexiaandClaude Opus 5
|
fe8f9d7457
|
[Docs] Add --prerelease=allow to cookbook uv install commands (#35920)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-21 15:51:48 -07:00 |
|
Xinyuan Tong
|
05c584c44f
|
docs: add DSPARK speculative decoding option to Ling-3.0-flash cookbook (#35861)
|
2026-08-22 02:50:55 +08:00 |
|
Jimmy Shong
|
3efa057449
|
[docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c (#35786)
|
2026-08-20 21:54:51 -07:00 |
|
Jimmy Shong
|
14795dcb1a
|
[docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev image tag (#35767)
|
2026-08-20 23:43:55 +00:00 |
|
Jimmy Shong
|
1a138e13b9
|
[docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753)
|
2026-08-20 23:34:44 +00:00 |
|
Jimmy Shong
|
d9f6861359
|
[docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35663)
|
2026-08-20 13:26:55 -07:00 |
|
Jason Wiemels
|
defb2a3100
|
feat(openai): Accept the input_audio content part in chat completions (#33606)
|
2026-08-19 13:37:50 -07:00 |
|
Xinyuan Tong
|
157d8ad27a
|
Support Intern-S2-Mobius FP8 (#34908)
|
2026-08-19 10:58:01 -07:00 |
|
 MickandClaude Opus 5
|
23f2320c95
|
[Docs] PaddleOCR-VL: update which stage of the pipeline this serves and show real output (#35458)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-19 21:27:58 +08:00 |
|
 MickandClaude Opus 5
|
77fc5c128e
|
[perf] overlap page preprocessing, pack the vit, enable prefill CUDA graph for paddle-ocr (#35318)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-08-19 08:20:55 +08:00 |
|