Commit Graph
18707 Commits
Author SHA1 Message Date
Mick 6c6175fabd perf: avoid excessive prefill CUDA graph padding (#31487) 2026-07-18 16:25:30 +08:00
Baizhou Zhang 38b29dcd6c Fix SM120 NVFP4 KV cache test OOM (#31653) 2026-07-18 00:53:47 -07:00
twb1235andZhiqiang Xie 071e649288 fix(rpc) Synchronize RPC requests only within the TP group. (#25213)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-07-18 07:44:51 +00:00
ThanhhaoandHao Phan 72c4ed1a3f [Spec] DFlash: remove per-step host syncs so the CPU runs a full step ahead (spec-v2 overlap) (#31468)
Co-authored-by: Hao Phan <htphan@nvidia.com>
2026-07-17 23:22:07 -07:00
Michael 7fbe91c6ea [AMD] register 8 JIT kernel benchmarks to jit-kernel-benchmark-test-amd (#31492) 2026-07-17 23:00:00 -07:00
Khoa PhamandClaude Opus 4.8 7a896215e7 [CP] Migrate MLA prefill CP (DeepSeek V3) to CP-v2 zigzag strategy (#31619)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 22:37:53 -07:00
Liangsheng Yin 639261f7b2 [Refactor] Move output logprob processing into the logprob_processor layer (#31624) 2026-07-17 22:33:52 -07:00
Liangsheng Yin 19c53c44a0 [Fix] Account zero-logprob sequences correctly in chunked logprob stitching (#31639) 2026-07-17 22:33:27 -07:00
DevashishLal-CBandDevashish Lal ab3d421c30 [plugin] oot torch profiler activity support (#31580)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
2026-07-17 22:20:54 -07:00
Alison Shao a5c0b94034 Let CI server launches wait longer for ports held by a dying predecessor (#31281) 2026-07-17 21:45:24 -07:00
44e4999ab2 [Diffusion] Use SGLang server for ERNIE-Image prompt enhancement (#31354)
Co-authored-by: Elizaveta Martirosian <you@example.com>
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-18 06:55:28 +03:00
R0CKSTAR 87dc211b87 [MUSA] Fix sglang-kernel build (#31634)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2026-07-17 20:51:09 -07:00
xdtbynd 359009fa00 [Bugfix][NPU] Fix/Refactor routed scaling factor application in MoE routing (#31449) 2026-07-18 10:59:13 +08:00
67e7f8d13a [JIT] Refactor dtype traits into DTypeTrait and unify warp reductions (#30838)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: jessiewei7 <jessiewei747@gmail.com>
Co-authored-by: root <root@GPUC5A6.maas>
2026-07-18 10:07:18 +08:00
Alison Shao e48eabbeee ci: add MLX to coverage report backend display order (#31442) 2026-07-17 17:32:20 -07:00
Brayden Zhong 238b2b2c9c Remove QServe and FBGEMM FP8 quantization (#31109) 2026-07-17 17:10:34 -07:00
Alison Shao f926c30c57 Revert "Fix mamba track-boundary seqlen under overlap scheduler (#31369)" (#31622) 2026-07-17 17:03:10 -07:00
Mick 42a058c760 optimize: avoid fla l2-norm recompilation by token count (#31558) 2026-07-18 07:58:36 +08:00
Qiaolin Yu 01a96720c6 [spec decoding] replace torch.multinomial with several native torch op in rejection sampling (#31620) 2026-07-17 16:58:11 -07:00
Baizhou Zhang 304a529558 Revert "Bump FlashInfer to 0.6.15 and revert regressions" (#31625) 2026-07-17 16:46:33 -07:00
Douglas YangandClaude Opus 4.8 a01a8e1ed9 docs(cookbook): replace pinned nightly/dev images with :latest (#31610)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 23:33:19 +00:00
Lianmin Zheng c95026aed3 Upgrade llguidance to 1.7.6 (#31484) 2026-07-17 16:31:44 -07:00
Qiaolin Yu 632adff9fd refactor logprob processor layer (#20071) 2026-07-17 16:28:10 -07:00
sglang-botandsglang-bot 0ad0ff2e9e chore: bump sglang-kernel version to 0.4.5 (#31618)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-17 15:58:36 -07:00
Qiaolin Yu e2d2e8d07e [spec decoding] fix multi_layer_eagle rotate_input_ids kernel registration (#31614) 2026-07-17 15:18:01 -07:00
Sam (Kesen Li) ec6a3163b7 [Feature] Add FP4 KV Cache Design and support SM120 GPUs (#21601) 2026-07-17 14:49:43 -07:00
Brayden ZhongandBrayden Zhong 7fc3fb9657 Remove deprecated Mamba flags from doc, wrong FP8 GEMM docstrings and change Nemotron image to 0.5.15 (#31094)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-17 14:34:29 -07:00
sglang-bot c00206c68c chore: bump sgl-kernel version to 0.4.5 (#31496) 2026-07-17 14:00:55 -07:00
Douglas YangandClaude Opus 4.8 ae3f62613a docs(cookbook): fix stale/pruned Docker image tags (#31508)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 13:24:46 -07:00
Liangsheng Yin 2c856abbe3 [Fix] Enable chunked input-logprob processing by default to cap peak memory (#31498) 2026-07-17 13:19:03 -07:00
Andy Ye d389039337 [diffusion] Opt in Qwen and Wan multi-output conditioning expansion (#31233) 2026-07-17 12:00:23 -07:00
fanxingranandkk fec6131844 [AMD] Disable DSA fused top-k v2 on ROCm for GLM-5.x / DeepSeek-V3.2 (#30506)
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
2026-07-17 10:22:14 -07:00
Muhammad Safiullah MemonandMuhammad Safiullah 2db21daf8d Fix Heartbeat Checker in KV Manager Disaggregation (#31584)
Co-authored-by: Muhammad Safiullah <muhammadsafiullah136@gmail.com>
2026-07-18 01:21:50 +08:00
Zhaoyi LiandMichael 2c64b7782e [AMD][PD] Fix early-send cached-prefix KV racing the prefill forward on mori (#31368)
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
2026-07-17 10:08:04 -07:00
Bingxu Chen 53229e88da [AMD] Fix stale imports in test_fused_fp8_kv_write.py (#31515) 2026-07-17 09:03:44 -07:00
Michael c546afc147 [AMD] Register 2 CPU/triton unit + kernel tests for AMD 1-GPU PR CI (#31379) 2026-07-17 09:01:46 -07:00
NOOBandR0CKSTAR 5e7eed4c00 [MLX] Honor --max-running-requests in the model runner stub (#30547)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-07-17 08:24:00 -07:00
Mick 85ac56c823 docs: simplify diffusion new model guide (#30109) 2026-07-17 19:39:38 +08:00
Mick 681c223570 refactor: wrap split backends once on full-attention backends (#31439) 2026-07-17 19:15:04 +08:00
Mick 24a8944e15 fix: enable Kimi multimodal breakable prefill cuda graph replay (#31391) 2026-07-17 19:13:54 +08:00
132ade55cd [Kernel] Rewrite JIT custom all-reduce (v2) with a decoupled kernel/storage design (#31049)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: root <root@GPUC5A6.maas>
2026-07-17 18:37:22 +08:00
Baizhou Zhang eaeb779ea4 [Doc] Update GLM5.2 Cookbook with LayerSplit usage (#31577) 2026-07-17 02:25:50 -07:00
Liangsheng Yin 19f4859b30 [CI] Exclude current process memory from GPU idle check (#31571) 2026-07-17 02:11:41 -07:00
zijiexiaandClaude Fable 5 8f765bc1c9 [Docs] Inkling cookbook: mark B300/GB300 recipes verified, tune B300 MTP mem fractions (#31550)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 16:07:09 +08:00
Mohammad Miadh AngkadandBrayden Zhong d67aa05697 Bump FlashInfer to 0.6.15 and revert regressions (#31502)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-07-17 00:50:12 -07:00
Baizhou Zhang e835512303 Add nightly test for GLM5.2 LayerSplit (#31512) 2026-07-17 00:29:14 -07:00
Xiaoyu ZhangandClaude Opus 4.8 619609aa5a [Kernel] Simplify sglang.kernels tests to idiomatic pytest style (RFC #29630) (#31546)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 15:06:26 +08:00
huangtingweiandZhangheng 44e3dd2713 [HiCache] Optimize HiCache host pool free-list release (#30658)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-07-17 14:52:08 +08:00
Liangsheng Yin 6e3be088a9 [Spec] Allocate the verify tree-mask scratch on the target backend only (#31527) 2026-07-16 23:23:21 -07:00
Junlin Wu bbd2a3fe4a ✨ [llm][npu][quant] Add W4A4 MXFP4 quantization support for Qwen3 Dense on Ascend NPU (#23795) 2026-07-17 09:06:30 +03:00