Commit Graph
1704 Commits
Author SHA1 Message Date
Chi McIsaacandMick Qian 6215aecd51 [diffusion] feat: add metrics support (#19084)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-18 13:22:03 +08:00
Khoa Pham b98a2d1096 [Docs] GLM-5.3-Flash cookbook: temporarily remove the DCP option (#40036) 2026-09-17 16:17:25 -07:00
Arseniy Mironov a1b4ec02ae [NPU][Diffusion] FA MXFP8 and modelslim w4a4f8 and w8a8f8 support for Wan2.2 and FLUX (#39438) 2026-09-17 13:13:05 +03:00
yl3469andShuwen Wang 1a90ae6727 Add Agentic-Aware Tail-Optimized LRU eviction to the unified radix cache (#34012)
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
2026-09-17 17:52:59 +08:00
faceless voidandronnie_zheng 44bd359082 [NPU][Diffusion] Optimize SenseNova-U1 batched generation (#39382)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-09-17 10:08:20 +03:00
MickandMick Qian c525ed8f02 [diffusion] fix: preserve explicit attention backends during autotune (#39882)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-17 13:28:55 +08:00
Liangsheng YinandBBuf 464fffbec8 dsv4.1: chat encoding and tool parsing (#39665)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-16 20:49:42 -07:00
Thomas Wang f60652a43e [AMD] Update deepseek-v4 PDI and cache policy setting for agentic workload (#39702) 2026-09-15 21:16:30 -07:00
jacky.chengandChangLiu0709 a9bb4d7d45 [AMD] Align Qwen3.5 MI355X HiCache cookbook with kernel / page_first (#39572)
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
2026-09-16 11:22:28 +08:00
amote-i 575163ff32 [NPU] [DOC] delete unsupported models in npu docs (#39555) 2026-09-15 14:55:17 +08:00
amote-i bdf8886ad3 [NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389) 2026-09-15 10:33:14 +08:00
a25f213bc4 [diffusion] docs: give the RTX 5090 its own H3 recipe, measured on a physical desktop (#39373)
Co-authored-by: Mick Qian <mickqian@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 09:15:13 +08:00
ChangLiu0709 242d8a70c0 [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260913, enable TOPK_V2 (#39406) 2026-09-14 21:50:52 +08:00
Theresa Shan 95140a7b0c [docs] DeepSeek-V4: MI355X PD disaggregation recipes for all three strategies (#39396) 2026-09-14 02:11:51 -07:00
jacky.cheng 2f5cc8e33e [AMD] Align Qwen3.5 MI355X cookbook with AttnFP8-V2 and HiCache direct / page_first_direct (#39358) 2026-09-14 12:46:17 +08:00
42b5af8c62 [diffusion] docs: refresh the VDN-H3 on b200 numbers in cookbook (#39244)
Co-authored-by: haochengxi <xihc@berkeley.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
2026-09-14 11:50:40 +08:00
6220f45d8e [PD] Add /v1/responses support to the HTTP PD router (#36141)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-13 22:48:13 +08:00
Shangming CaiandXinyuan Tong 14a131ad5b [PD][OpenAI] Gate /v1/responses persistence behind --enable-response-store, default off (#39122)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-13 21:20:47 +08:00
cebca698e2 [Qwen3.8] Enable NVIDIA NVFP4 on DGX Spark with file-backed PLE and PDL router fix (#39126)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: rdxa <rdxa@rdxa-int-spark-01.yvb.moe>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Manrique <nanomlm@gmail.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
2026-09-13 16:23:41 +08:00
Thomas Wang ec5fba5777 [AMD] Add dspark config and agentic workload section for deepseek-v4 model (#39252) 2026-09-12 19:32:01 -07:00
Xiaoyu ZhangandMick Qian 23bc4c6ed9 [Diffusion] Return Qwen-Image-Layered outputs and preserve CFG2 rounding (#38549)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-13 09:15:47 +08:00
Zhang, Jiejing 288627e400 [AMD] Document GLM-5.2 MXFP4 recipe update on MI355X (#39230) 2026-09-12 15:57:15 -07:00
Xinyuan Tong b5a2aebc7e [Docs] GLM-5.3-Flash cookbook: fixed MTP 5/1/6, EP1 + flashinfer_trtllm on Blackwell (#39213) 2026-09-12 12:56:57 -07:00
desmond-intel 6dc7b3421b Inference Support Mamba 2 and 1 (#34556) 2026-09-12 20:51:18 +08:00
ChangLiu0709 7bc4eb3740 [AMD] Update MI355X MXFP4 HiCache defaults and quick-reduce quantization for Qwen3.5 cookbook (#39104) 2026-09-12 00:59:32 -07:00
Khoa PhamandClaude Fable 5.1 6ba96d329f [DCP] Resolve --dcp-comm-backend to fi_a2a/a2a by default for every model (#39165)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 23:48:42 -07:00
Khoa PhamandClaude Fable 5.1 0d08668821 [Cookbook] Kimi-K3: keep DCP under HiCache L1+L2 with DSPARK (#39190)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 23:14:24 -07:00
ff1ce11348 [diffusion] model: support VDN-H3 with a hybrid_window_attn_h3 backend (#37903)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Haocheng Xi <xihc@berkeley.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-09-12 11:36:32 +08:00
Vignesh Sethuraman 0d1bea77da [AMD] Allow aiter attention backend for Gemma-4 (#38758) 2026-09-11 18:48:26 -07:00
Kevin Mi 5f3606c7b2 [Cookbook][AMD] Kimi-K3 MI350X/MI355X: pin a ROCm image with the DSPARK graph-capture fix, add measured cell numbers (#39029) 2026-09-11 16:10:44 -07:00
df6424967a [docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Jiminator <jimmysh341@gmail.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
2026-09-11 11:58:30 -07:00
Zhang, Jiejing d7c284b894 [AMD] Use the triton DSA backend for GLM-5.2 MXFP4 on MI355X (#39106) 2026-09-11 11:48:50 -07:00
MickandMick Qian e016de462c [diffusion] CI: expose nightly server telemetry coverage (#38782)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-11 23:07:37 +08:00
Siju Samuel 67d3a2ea57 [XPU] Make checkpoint_engine worker device-agnostic (#32382) 2026-09-11 09:56:39 +08:00
Mohammad Miadh AngkadandMohammad Angkad 52c191da52 [Deps] Retire the CUDA 12 lane (#38404)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-10 16:58:09 -07:00
sglang-botandsglang-bot 887c401e15 docs: sync LMSYS SGLang blog cards (#36773)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-09-10 18:37:14 +00:00
zijiexiaandClaude Opus 5 4b7331fb77 Make the remaining DeepSeek-V4.1 NVIDIA cells start (#38861)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 11:04:53 -07:00
Yuhao YangandClaude Code a37ded1693 [Cookbook] DeepSeek-V4.1: add the HiCache L2 knob to the Playground (#38844)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-10 18:20:34 +08:00
MickandMick Qian c9c26d56b2 [diffusion] docs: sync snapshot and minimax-h3 subblock features (#38784)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-10 17:23:43 +08:00
zijiexiaandClaude Opus 5 5caafd2118 Fix the DeepSeek-V4.1 reasoning example and mark the B300 cells verified (#38839)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 01:49:44 -07:00
Xinyuan Tong 9a2f17f41d Add INT4 and FP4 lanes to the Ling-3.0-flash-VL cookbook (#38527) 2026-09-10 15:23:50 +08:00
1b77f498a0 [NVIDIA] Support flashinfer Mega Moe (#31470)
Co-authored-by: djns99 <40156487+djns99@users.noreply.github.com>
Co-authored-by: 云挚 <ningyunxiao.nyx@antgroup.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-10 00:22:47 -07:00
Brayden ZhongandBrayden Zhong c0b790cf7f Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-09-10 15:12:01 +08:00
zijiexiaandClaude Opus 5 69777c4d36 Add DeepSeek-V4.1 Flash cookbook (#38802)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 23:15:30 -07:00
Thomas Wang c415f977b8 [AMD] Update v4 args for agentic workload (#38677) 2026-09-09 21:41:05 -07:00
MickandMick Qian ce555ed82a [diffusion] refactor: refactor utility ownership and document helper placement (#38699)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-10 09:11:01 +08:00
Alison Shao a7e00b7576 [CI] Answer unrecognized slash commands instead of skipping silently (#38736) 2026-09-09 18:01:18 -07:00
William Hu 0084030179 Add Opt-In for GLM-5.3 Flash breakable prefill CUDA graphs (#38522) 2026-09-09 17:02:08 -07:00
Alison Shao 2948a62a6f [CI] Add /run-full-ci and /run-extra-ci slash commands (#38734) 2026-09-09 13:56:38 -07:00
Even Zhou dba34cc964 [NPU] Bump memfabric and sgl-kernel-npu versions in docs and pyproject_npu.toml (#38437) 2026-09-09 19:53:49 +08:00