5089 Commits
Author SHA1 Message Date
KnightYao 7ef49e8f7d Rename Spark3 to Spark2.5 (#36416) 2026-08-26 00:04:46 -07:00
Ziang Li 3c9febc68b [Spec][DSA] Add --speculative-dsa-topk-backend (#36313) 2026-08-25 23:35:03 -07:00
Tan TrinhandKhoa Pham 04c1036bb3 [Kernel] Skip reserved writes in MLA KV cache (#36003)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-08-25 22:42:23 -07:00
34de1fb47f fix(test): stabilize nightly precision regression (#34668)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-08-25 20:04:52 -07:00
HuangJi cc3b61873f [Diffusion][minimax-h3] Restrict MiniMax-H3 SubBlock sparsity to video queries (#35850) 2026-08-26 10:55:09 +08:00
Alison Shao 4382947b58 Fix DSV4 shared-fusion CPU unit test after the EP guard landed (#36424) 2026-08-25 19:52:05 -07:00
Shuwen WangandClaude Opus 5 303227951a fix: use a bf16-relative tolerance in the DSA indexer K kernel test (#35795)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 19:50:29 -07:00
Bingxu Chen 07a9de25b4 [AMD][CI] Fix shared-KV verify tests for multi-head GQA (#36296) 2026-08-25 19:46:27 -07:00
Yuanle Liu 4ae30dc736 fix: detect cross-node multimodal transport by nnodes (#35646) 2026-08-26 10:31:48 +08:00
2d8484740d Support deepseek v4 and kimi k3 on ssd (#35314)
Co-authored-by: 1BIN4 <1741738350@qq.com>
Co-authored-by: L-Ark <fliangae@connect.ust.hk>
Co-authored-by: Chikati <jxudn@connect.ust.hk>
Co-authored-by: mengzili <zilim@ust.hk>
2026-08-26 10:05:12 +08:00
Zhangheng bec6248272 [Unified Cache]: add glm5.2 per commit ci (#36281) 2026-08-26 10:04:19 +08:00
Mick 223dfce917 fix: bound CUDA memory for fast image preprocessing (#36295) 2026-08-26 09:02:08 +08:00
Baizhou ZhangandRyan Stewart 41e7612dee [Model] Support Nemotron 3.5 Lightning speculative decoding (#36186)
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-25 16:43:58 -07:00
Baizhou Zhang 2d88c79b3e [CI] Add Kimi-K3 MMMU-Pro accuracy coverage (#36284) 2026-08-25 16:33:46 -07:00
cctry aa718f7343 Refactor HiCache host pool management (#36232) 2026-08-25 16:31:57 -07:00
0c42a44cd7 Add Spark3 Model (#35963)
Co-authored-by: Yaowj <yaowj@MacBook-Air.local>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
2026-08-25 16:24:00 -07:00
Leon Gao f7a56494b1 Fix SWA ownership across grouped frees (#36381) 2026-08-25 14:57:30 -07:00
Junpan Wu 4b4bf3d2a5 [Deepseek-V4] Enable shared-experts fusion on the flashinfer_mxfp4 (trtllm-gen) MoE path (#35505)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
2026-08-25 14:55:25 -07:00
Michael 6569125e3a [AMD][CI] Add MiniMax-M3-MXFP8 MI35x nightly perf benchmark (#36142) 2026-08-25 14:05:01 -07:00
Kevin FlansburgandShangming Cai 7ddf92d5f4 fix(disagg): refresh stale prefill bootstrap metadata (#36029)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-08-26 03:29:20 +08:00
Bingxu Chen 96a73c4e15 [AMD][CI] Adjust MI300 score API performance thresholds (#36290) 2026-08-25 09:41:36 -07:00
YAMY e9c9df6a52 [Performance] Tune FlashInfer EXTEND for DP prefill (#36219) 2026-08-25 08:29:57 -07:00
elvischenv 46d9427b91 Fix MXFP8 MoE weight sizing for non-gated models (#36097) 2026-08-25 22:09:09 +08:00
829138a31e [HiCache] Fix PP inconsistency with HiCache L3 (#22607) (#27010)
Co-authored-by: ybyang <ybyang7@iflytek.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: shangmingc <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-08-25 20:49:57 +08:00
Cheng Wan 443527af0c config: the model-config cache keys on the path the record carried (#36300) 2026-08-25 03:18:28 -07:00
Bingxu Chen d067622820 [AMD][bugfix] Add moe_ep_size/moe_tp_size to the allreduce-fusion gate test stub (#35340) 2026-08-25 01:24:52 -07:00
Lianmin Zhengandyangliu991 bf1e03f712 [MegaMoE] Respect padded MXFP8 scale row strides in pre-dispatch (#36237)
Co-authored-by: yangliu991 <yangliu991@fb.com>
2026-08-25 00:27:33 -07:00
e2b50930b9 [NPU] DeepSeek-V4 adapt sgl-kernel-npu ops (compressor/sparse-attn/sparse-attn-metadata) (#35676)
Co-authored-by: vstone-w <374330057@qq.com>
Co-authored-by: unclezhou486 <154310456@qq.com>
Co-authored-by: 摆渡人 <2044145178@qq.com>
2026-08-25 15:13:12 +08:00
MingxuZhandClaude b7f9fca26e [CPU] Raise mem-fraction-static 0.1->0.2 in intel_amx backend a/b (#36260)
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-25 14:36:20 +08:00
chuyehandchuyeh 13d2ee8180 [AMD] Skip shared-KV verify test on ROCm 7.0 CI (#36139)
Co-authored-by: chuyeh <chuyeh@users.noreply.github.com>
2026-08-24 22:20:24 -07:00
Baizhou Zhang 833be86c15 [CP V1 Deprecation 1/5] Migrate tests to strategy-based prefill CP (#36222) 2026-08-24 20:15:34 -07:00
longxin9715 a0f52eca98 [NPU] Add test for --dllm-fdfo (#33634) 2026-08-25 10:58:45 +08:00
Alex NailsandClaude Opus 5 998eeda0a5 [CI] Register test_dflash_logits at its real cost (#36242)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 18:22:49 -07:00
Alex NailsandClaude Opus 5 5ffdb02d0c [CI] Cut repeated tokenizer loads, serial subprocesses and a double scan (#36241)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 18:22:43 -07:00
Alex NailsandClaude Opus 5 68575b23d0 [CI] Stop the config ratchets re-parsing the package on every scan (#36240)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 18:22:32 -07:00
pllimax d2d8ecea77 [npu] Combine NPU test fixes from #35472 and #34516 (#36180) 2026-08-25 09:00:40 +08:00
huangtingwei 0f7ba3d115 [HiSparse] Support hisparse multi-step swap io kernel (#32162) 2026-08-24 17:29:44 -07:00
Alex NailsandClaude Opus 5 1ec20fd25d [CI] Stop the resolution ratchets re-parsing the whole package (#36235)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 16:55:35 -07:00
effe0d14d2 [Fix] Keep the MiniCPM-SALA config reads visible to the resolution ratchets (#36178)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-24 16:05:09 -07:00
BingjiaWang 3c481b9421 [Benchmark] Add optional steady-state window for serving metrics (#30918) 2026-08-24 14:49:20 -07:00
AMD-yanfeiwang 24bce93c93 [AMD][MORI] Deduplicate CP-replicated state transfers (#36025) 2026-08-24 14:43:52 -07:00
jacky.cheng 0665740953 [AMD][Fix] Route MoRI through the Qwen MoE all-to-all path (#32039) 2026-08-24 13:04:08 -07:00
Ke Bao 586211bc46 Add PD test for inkling with mxfp8 KV (#35840) 2026-08-25 00:12:05 +08:00
Ke Bao 54ec2c4699 Fix recurrent state loss on decode retraction (#35957) 2026-08-25 00:11:05 +08:00
fzyzcjy e586a6f2c5 Report the whole server's world size in the scheduler's internal state (#35929) 2026-08-24 20:21:45 +08:00
fzyzcjy 6dd79576cd Expose the declared sglang env vars of a scheduler in its internal state (#35928) 2026-08-24 20:20:29 +08:00
fzyzcjy c56cee0f80 Support gated launch to defer startup memory allocation (#35927) 2026-08-24 20:19:45 +08:00
fzyzcjy 3b24d8981b Report per-token weight-version spans in generation meta info (#35926) 2026-08-24 20:18:52 +08:00
fzyzcjy 981dfa2b83 Make the scheduler track the published weight version (#35925) 2026-08-24 20:17:50 +08:00
Hert4andAlex Nails b4bd5f91ee [diffusion] Fix test_model_fast_paths import after sana_ln_modulate rename (#36175)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-24 04:48:17 -07:00