Commit Graph
1680 Commits
Author SHA1 Message Date
ChangLiu0709 7bc4eb3740 [AMD] Update MI355X MXFP4 HiCache defaults and quick-reduce quantization for Qwen3.5 cookbook (#39104) 2026-09-12 00:59:32 -07:00
Khoa PhamandClaude Fable 5.1 6ba96d329f [DCP] Resolve --dcp-comm-backend to fi_a2a/a2a by default for every model (#39165)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 23:48:42 -07:00
Khoa PhamandClaude Fable 5.1 0d08668821 [Cookbook] Kimi-K3: keep DCP under HiCache L1+L2 with DSPARK (#39190)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 23:14:24 -07:00
ff1ce11348 [diffusion] model: support VDN-H3 with a hybrid_window_attn_h3 backend (#37903)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Haocheng Xi <xihc@berkeley.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-09-12 11:36:32 +08:00
Vignesh Sethuraman 0d1bea77da [AMD] Allow aiter attention backend for Gemma-4 (#38758) 2026-09-11 18:48:26 -07:00
Kevin Mi 5f3606c7b2 [Cookbook][AMD] Kimi-K3 MI350X/MI355X: pin a ROCm image with the DSPARK graph-capture fix, add measured cell numbers (#39029) 2026-09-11 16:10:44 -07:00
df6424967a [docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Jiminator <jimmysh341@gmail.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
2026-09-11 11:58:30 -07:00
Zhang, Jiejing d7c284b894 [AMD] Use the triton DSA backend for GLM-5.2 MXFP4 on MI355X (#39106) 2026-09-11 11:48:50 -07:00
MickandMick Qian e016de462c [diffusion] CI: expose nightly server telemetry coverage (#38782)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-11 23:07:37 +08:00
Siju Samuel 67d3a2ea57 [XPU] Make checkpoint_engine worker device-agnostic (#32382) 2026-09-11 09:56:39 +08:00
Mohammad Miadh AngkadandMohammad Angkad 52c191da52 [Deps] Retire the CUDA 12 lane (#38404)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-10 16:58:09 -07:00
sglang-botandsglang-bot 887c401e15 docs: sync LMSYS SGLang blog cards (#36773)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-09-10 18:37:14 +00:00
zijiexiaandClaude Opus 5 4b7331fb77 Make the remaining DeepSeek-V4.1 NVIDIA cells start (#38861)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 11:04:53 -07:00
Yuhao YangandClaude Code a37ded1693 [Cookbook] DeepSeek-V4.1: add the HiCache L2 knob to the Playground (#38844)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-10 18:20:34 +08:00
MickandMick Qian c9c26d56b2 [diffusion] docs: sync snapshot and minimax-h3 subblock features (#38784)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-10 17:23:43 +08:00
zijiexiaandClaude Opus 5 5caafd2118 Fix the DeepSeek-V4.1 reasoning example and mark the B300 cells verified (#38839)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 01:49:44 -07:00
Xinyuan Tong 9a2f17f41d Add INT4 and FP4 lanes to the Ling-3.0-flash-VL cookbook (#38527) 2026-09-10 15:23:50 +08:00
1b77f498a0 [NVIDIA] Support flashinfer Mega Moe (#31470)
Co-authored-by: djns99 <40156487+djns99@users.noreply.github.com>
Co-authored-by: 云挚 <ningyunxiao.nyx@antgroup.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-10 00:22:47 -07:00
Brayden ZhongandBrayden Zhong c0b790cf7f Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-09-10 15:12:01 +08:00
zijiexiaandClaude Opus 5 69777c4d36 Add DeepSeek-V4.1 Flash cookbook (#38802)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:15:30 -07:00
Thomas Wang c415f977b8 [AMD] Update v4 args for agentic workload (#38677) 2026-09-09 21:41:05 -07:00
MickandMick Qian ce555ed82a [diffusion] refactor: refactor utility ownership and document helper placement (#38699)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-10 09:11:01 +08:00
Alison Shao a7e00b7576 [CI] Answer unrecognized slash commands instead of skipping silently (#38736) 2026-09-09 18:01:18 -07:00
William Hu 0084030179 Add Opt-In for GLM-5.3 Flash breakable prefill CUDA graphs (#38522) 2026-09-09 17:02:08 -07:00
Alison Shao 2948a62a6f [CI] Add /run-full-ci and /run-extra-ci slash commands (#38734) 2026-09-09 13:56:38 -07:00
Even Zhou dba34cc964 [NPU] Bump memfabric and sgl-kernel-npu versions in docs and pyproject_npu.toml (#38437) 2026-09-09 19:53:49 +08:00
daf66f6670 [Diffusion][SenseNova] support SenseNova-U1.5-8B-MoT (#36606)
Co-authored-by: wuyuefeng <wuyuefeng@noreply.gitcode.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-09-09 10:12:38 +03:00
Xiaoyu Zhang a27be5ff62 [Diffusion] Enable lossless BCG for FLUX.1-dev (#38591) 2026-09-09 14:09:21 +08:00
MickandMick Qian 00a9028e87 [diffusion] feat: add explicit snapshot-offload component residency (#38535)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-09 13:44:54 +08:00
Xiaoyu Zhang 03d06a764e [Docs] Add measured JoyEcho H200 residency and BCG recipe (#38534) 2026-09-09 11:19:27 +08:00
Baizhou Zhang e54ff1efb9 [CP V1 Deprecation 5/5] Update prefill CP documentation (#36230) 2026-09-08 19:27:49 -07:00
MickandMick Qian a8e45f16cc [diffusion] feat: support mixed INT8 embeddings and Comfy NVFP4 encoders for minimax-h3 (#38506)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-09 08:36:43 +08:00
MickandMick Qian 65400bb420 [diffusion] model: support MiniMax-H3 singularity hybrid checkpoints (#38455)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-09 08:35:39 +08:00
Cheng Wan db272201a2 [Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement (#38375) 2026-09-08 16:42:12 -07:00
Xinyuan Tong afe90a8bc9 Point Ling-3.0-flash-VL cookbook install section at the model image (#38539) 2026-09-08 10:51:33 -07:00
Xinyuan Tong 482e9f257b Add Ling-3.0-flash-VL cookbook (#38434) 2026-09-08 23:00:24 +08:00
HZY a6b542813f fix(glm-5.2-nvfp4): bound Mooncake synchronous transfer batches (#32758) 2026-09-08 22:14:33 +08:00
Wuhen Duan dfd9b5c2a4 [NPU] Enable non-greedy MTP sampling (#32495) 2026-09-08 11:18:01 +08:00
fba967ed9c [diffusion] Support diffusion decoder parallel tiling for LTX-2.5 (#36026)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-08 10:01:06 +08:00
Siju Samuel 2358916d5a [Feature][Intel XPU] Add memory saver support for Intel XPU via upstream torch_memory_saver (#29935) 2026-09-08 09:38:51 +08:00
Kedar PotdarandPo-Han Huang f4bbf12423 docs(cookbook): Qwen3.5 FP8 on B200/B300 — trtllm-gen MoE + symm mem (#38374)
Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
2026-09-08 09:12:49 +08:00
f4b75b5c36 docs(cookbook): Qwen3.8-Flash-Next NVFP4 recipes for DGX Spark (1x, 2x) and RTX PRO 6000 (#37995)
Co-authored-by: Jiminator <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-07 15:33:03 -07:00
zijiexiaandClaude Opus 5 e4008de757 Add MiniCPM5-2B cookbook (#38295)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 21:30:17 +08:00
Mick ba6d3df69a [diffusion] doc: document verified GB300 and derived GB200 H3 recipes (#38296) 2026-09-07 16:32:29 +08:00
Mick 15d2cbcc90 [diffusion] CI: validate every repeated server request (#38185) 2026-09-07 14:36:50 +08:00
39a80354aa [MUSA] Add installation guide and Dockerfile (#36709)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-09-06 20:13:53 -05:00
faceless voidandgithub-actions[bot] 30d0eb2ca9 [NPU] Adapt DFlash2 speculative decoding to Ascend NPUs (#35629)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-07 09:08:39 +08:00
Mick f3d05644db [diffusion] docs+skill: document which components to stream under layerwise offload (#35674) 2026-09-06 23:12:36 +08:00
MickandClaude Fable 5 ade1da017f [diffusion] docs: verify the DGX Spark H3 recipe (#37456)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-06 23:10:27 +08:00
Mick 938dc5621d [diffusion] refactor: reuse plain state-dict loading without per-model classes (#38127) 2026-09-06 18:39:21 +08:00