Commit Graph
17588 Commits
Author SHA1 Message Date
Liangsheng Yin 83a9b5dd88 [mem_cache] Drop the torch.unique sync from the SWA page expansion (#37463) 2026-09-01 14:16:47 -07:00
b24c8f10e7 [FlashInfer] Avoid D2H sync for sliding-window lengths (#32218)
Co-authored-by: llilian73 <204300658+llilian73@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-09-01 14:08:22 -07:00
zijiexiaandClaude Fable 5 0f18d389b4 [Cookbook] Verify DeepSeek-V4 Flash Vision balanced and high-throughput on B200 (#37468)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 13:09:14 -07:00
Xinyuan Tong 442c7c1e29 [Docs] GLM-5.3-Flash cookbook: add NVFP4 FP8+TRT-LLM benchmark rows (follow-up to #37109) (#37412) 2026-09-01 12:56:22 -07:00
Cheng Wan 0b1ce3d140 [Feature] Unified memory: support decode context parallelism for Kimi-Linear (#36890) 2026-09-01 12:44:26 -07:00
Ankur Singh 3315356cc0 docs(cookbook): enable FlashInfer GDN for Qwen3.5 B200 (#37360) 2026-09-01 11:48:41 -07:00
Shuwen Wang a2b8681d1d [CI][MLX] Restore the mamba_branching_seqlen attribute the MLX runner reads off a request (#37453) 2026-09-01 11:16:45 -07:00
cctryandcctry 9a05b470fa [Memory] Size the CUDA graph pool from warmup measurements and fix graph-pool borrowing (#36911)
Co-authored-by: cctry <cctry@fb.com>
2026-09-01 09:32:38 -07:00
c34f378342 fix(nixl): make FILE path-mode devId globally unique (#34362)
Co-authored-by: hekh <hekh@local>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 09:11:59 -07:00
Thomas Wang bb3e3cbceb [AMD] Fix v4 topk issue (#37439) 2026-09-01 08:37:30 -07:00
billishyahao 44a92e54b9 [AMD] fix aiter cannot get heuristic kernel regression (#37438) 2026-09-01 08:35:31 -07:00
DarkSharpnessandClaude Opus 5 cb6dd58fbe [Kernel] Replace dsv3_router_gemm with the unified tiny GEMM (#34693)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 23:00:21 +08:00
ee462b5899 [Kernel] Add tuned LFM2.5 Triton MoE configs on B300 (#37158)
Co-authored-by: Song Bian <biansonghz@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 22:55:38 +08:00
Xiaoyu Zhang 5993f91f84 [Kernel] Register merged diffusion agent kernels with KDA backend (#37385) 2026-09-01 22:51:10 +08:00
Xiaoyu Zhang 4c7ff0d906 [CI] Double JIT kernel unit test timeout (#37435) 2026-09-01 22:37:37 +08:00
DarkSharpnessandClaude Opus 5 b6c06e1efb [DSA] Drop the redundant 512 from the top-k transform entry-point names (#36831)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 22:37:03 +08:00
Even Zhou 00689c0c94 [CI] Add entrypoint to hc_combine test for standalone execution (#37406) 2026-09-01 22:08:00 +08:00
3ae54c6ca2 test(npu): add DSV4-Flash / GLM-5.2 / Kimi-K3 gpqa accuracy cases (#37431)
Co-authored-by: Sugar920 <Sugar920@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-01 22:00:04 +08:00
Estrella-xx 4c2c169e6b [NPU]Strip padding before FIA kernel for vision encoder padded sequences (#36329) 2026-09-01 19:28:19 +08:00
YC Yen-Ching TsengandChen 3103bc7462 [AMD][CI] Add daily ROCm 10 and ROCm 7.2 test coverage (#37409)
Co-authored-by: Chen <bingxche@amd.com>
2026-09-01 19:23:19 +08:00
ce7e79b32c [AMD] Fix nightly ROCm 7.0 image build: patch missing <optional> include in AITER topk kernel (#36216)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Chen <bingxche@amd.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-09-01 19:06:32 +08:00
Eric.Chin.AMDandKingRei ed122ea984 [AMD] Enable topk v2 GLM ROCm (#36851)
Co-authored-by: KingRei <hiroki1139@gmail.com>
2026-09-01 03:48:50 -07:00
YC Yen-Ching Tseng b425897366 [AMD] Gate the aiter memory-reserve exemption behind an env var (#37242) 2026-09-01 03:46:15 -07:00
datdo-msftandShangming Cai 49db27528a fix(test): deflake zmq load-snapshot round-trip tests (#35787)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-09-01 18:14:55 +08:00
Jialin Ouyang a77283fb02 [Rust] Rename mem-cache to sglang-radix-tree (#37290) 2026-09-01 03:03:44 -07:00
chx96642264 c16a8fc899 [NPU] [bugfix] Fix NPU MLA HiCache backup accessing missing data_ptrs. (#36813) 2026-09-01 17:31:11 +08:00
zijiexiaandClaude Opus 5 dc1ae02684 [Cookbook] Add the DFlash2 speculative option to GLM-5.3 (#37392)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 09:08:28 +00:00
eeecho b68702be99 [DSV4] hc-prenorm: fuse the combine step into a Triton kernel (#35118) 2026-09-01 02:04:05 -07:00
zijiexia 6c72b49a57 Revert "[AMD] Add GLM-5.3-Flash recipes for MI300X, MI325X, and MI355X (#36608)" (#37380) 2026-09-01 01:25:13 -07:00
Ziang Li 5edcd0a445 [FlashInfer V0.6.18] feat(dsv4): support --dsa-topk-backend flashinfer with fused top-k (#33237) 2026-09-01 01:18:10 -07:00
Liangsheng Yin 3484f7f836 [mem_cache] Add free_kv_row to release a request's kv row by row range (#36721) 2026-09-01 01:14:43 -07:00
Xiaoyu ZhangandCursor 1c3ad92438 [Diffusion] Fuse FLUX.2 ModelOpt FP8 producers and QKV packing (#37162)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 16:14:29 +08:00
zijiexiaandClaude Fable 5 379e33d87e [Cookbook] Add NVFP4 options for DeepSeek-V4 Flash Official (0731) and Pro Official (0813) (#37351)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 08:03:39 +00:00
Mohammad Miadh Angkad 03b33cbe5d [CI] Fix hybrid wrapper test fake missing kv_index_translator (#37374) 2026-09-01 00:55:05 -07:00
chien-an-chen 8a191554e3 [AMD] Enable 12-head MLA aiter fp8 Gluon decode (batched bh16bn128). (#34647) 2026-08-31 23:47:45 -07:00
Ma Mingfei 5e79110122 [Fix][CPU] fix xeon ci failure by test_qwen35_flashinfer_fusion (#37338) 2026-09-01 14:17:35 +08:00
Xinyuan Tong 60548501bb [Docs] Add NVFP4 section to GLM-5.3-Flash cookbook (#37109) 2026-09-01 14:14:30 +08:00
Liangsheng Yin 959ca033eb refactor(hicache): simplify decode offload state bookkeeping (#37299) 2026-08-31 23:09:41 -07:00
Mick ae2bd5728b [vlm] fix: contain multimodal feature transport failures (#37047) 2026-09-01 13:46:38 +08:00
YAMY 33ed29a0ee test: update hybrid attention runner fixtures (#37345) 2026-08-31 22:00:50 -07:00
Zhaoyi Li 458987b5ac [AMD][MORI] Bump MoRI to 7c51d18 for ionic RoCE dmabuf fix (#509) (#37286) 2026-08-31 21:37:05 -07:00
huangtingweiandhzh0425 b21000aef1 [Unified Cache][4/N]: Add Mooncake backend for external linker (#37205)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-09-01 11:40:06 +08:00
Liangsheng Yin 3b14f37b74 [Fix] Use real ReqKvInfo in unit-test req mocks (#37339) 2026-08-31 20:25:51 -07:00
hhhh1252023 60f881b40c [CI/NPU] Isolate multi-node tests by run_id to prevent concurrent-run… (#35500) 2026-09-01 11:08:27 +08:00
Yuan Luoandluoyuan.luo 5b04408784 [MoE] Add FlashInfer SM90 MXFP4 W4A8 CUTLASS MoE (#34967)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-08-31 20:04:41 -07:00
22337e9c56 fix(unified-memory): forward the KV-index translator through every wrapper backend (#37307)
Co-authored-by: Caihua Li <caihua.li@bytedance.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-08-31 19:55:21 -07:00
Mohammad Miadh Angkad e6f21cdadc [Cohere Command-A-Plus] Optimize decode and BCG capture on SM10X (#36624) 2026-08-31 19:37:47 -07:00
Yiqi Yang 1591dcd91a [diffusion] fix: fix loading a block-FP8 quantized MiniMax-H3 DiT (#35703) 2026-09-01 10:33:58 +08:00
71cee04ebe [Diffusion] Optimize Qwen-Image TP collectives and attention (#36680)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:28:37 +08:00
Martin Hua 562b661e0e [Feature] Megatron LayerNorm sequence parallelism (--enable-layernorm-sp) (#30915) 2026-08-31 19:27:28 -07:00