Commit Graph
15801 Commits
Author SHA1 Message Date
Ke Bao 3763d3fa8c Add decode-lock skip to compute-mamba-ratio (#32789) 2026-07-29 22:38:34 +08:00
Wu Jiangming 1c6a0e91e1 fix mqa preshuffle layout issue for deepseek v4 (#31563) 2026-07-29 07:36:48 -07:00
Mick 22151edca1 [diffusion] optimization: accelerate CUDA video output finalization (#32784) 2026-07-29 22:04:20 +08:00
Xiaoyu Zhang 0ebbe43dbb fix(diffusion): size VSA top-k from padded blocks (#32695) 2026-07-29 21:58:41 +08:00
Xiaoyu Zhang 4f5b50c576 perf(diffusion): decode Wan VAE in BF16 (#32697) 2026-07-29 21:57:50 +08:00
Xiaoyu Zhang 917e900d4d feat(diffusion): add regional torch compile (#32696) 2026-07-29 21:57:05 +08:00
Ke Bao 50b029257f Skip mamba lock during decoding (#32228) 2026-07-29 21:53:28 +08:00
pllimax d004a15a3e Fix GLM4-7B-Flash accuracy test configuration, tune Qwen3.6-27B/35B performance test parameters, and harden Ascend NPU multi-node E2E test utilities against pod name format errors. (#32371) 2026-07-29 21:44:17 +08:00
Ilia Yastrebov 977f04aafe [PD] NIXL connector: shard by destination (#32025) 2026-07-29 21:02:13 +08:00
MickandClaude Sonnet 5 67c2258906 [diffusion] fix: fix dual-DiT models crash with (1,)-placeholder weights after compile-time offload (#32743)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 20:44:13 +08:00
JiaruiChang5268 ca6e0ff2c8 [NPU] fix dsv4 mtp condition on NPU graph (#32711) 2026-07-29 20:13:31 +08:00
Yi (Vincent) Zhong d254ec9ff8 Add LFM2.5 embedding model support (#28691) 2026-07-29 11:58:47 +00:00
LZW 4c82bb3252 Add Mooncake tenant id support (#30256) 2026-07-29 18:17:41 +08:00
Xiaoyu ZhangandClaude Fable 5 8742a1a0f8 Add a benchmark script for the HPC-Ops bf16xfp32 router GEMM (#32642)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:17:08 +08:00
Ke Bao c151080d28 Add compute-mamba-ratio skill (#32763) 2026-07-29 18:15:48 +08:00
Baizhou Zhang d12ea3e9ba docker: add Kimi K3 images (#32760) 2026-07-29 03:02:28 -07:00
hunhokimandHun-ho Kim 983e4aa18d Eliminate redundant DSA state transfers (Mooncake) (#32620)
Co-authored-by: Hun-ho Kim <hunho.kim@samsung.com>
2026-07-29 17:58:48 +08:00
Xinyu JiangandZhiyao Jiang f5bcd00e16 [AMD] DSv4: bring HIP compress-state pool into the memory_saver KV_CACHE region (#31747)
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
2026-07-29 02:41:08 -07:00
Xiaoyu Zhang c32c4ef79c [Kernel] Move sgl-kernel under sglang.kernels.aot (#32648) 2026-07-29 17:25:00 +08:00
Mohammad Miadh Angkad 1b9dfa14e6 Fix FlashInfer MNNVL workspace size check (#32318) 2026-07-29 02:12:18 -07:00
Junlin Wu 9ab88380c1 👥 chore(codeowners): Update codeowners for NPU quantization (#31917) 2026-07-29 12:10:01 +03:00
Yihao Wang 227dadd79a [diffusion] feat: support resident layers for DiT (#31538) 2026-07-29 16:52:55 +08:00
Andrew KuksaandANDREW_K 0caf0fc01d [Diffusion][Docs] Ascend A2, A3 add basic usage and benchmark results in diffusion cookbook (#30614)
Co-authored-by: ANDREW_K <andrewsha3@DESKTOP-KNDINTT.localdomain>
2026-07-29 11:50:15 +03:00
Brayden ZhongandBrayden Zhong 7dcebca255 Fix nightly CI: NVFP4 cuda-graph crash, NVILA batching, CuTe paged-KV zero-size, Kimi-VL OOM (#32118)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-29 01:39:55 -07:00
Liangsheng Yin cce5873513 [CI] Fail lint when a registered file's TestCase classes never run (#32735) 2026-07-29 00:45:01 -07:00
f05c92fb6d [llm][npu][quant] Add W8A8 MXFP8 quantization for Qwen3 MoE on Ascend NPU (#30768)
Co-authored-by: Артем Савкин <58187114+OrangeRedeng@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-29 10:39:36 +03:00
Mick da5528db30 fix(vlm): materialize Qwen3-VL features on the vision device (#31596) 2026-07-29 15:25:04 +08:00
weireweireandweireweire bd47ec97ff [EAGLE] Handle NaNs in fused top-k=1 (#32396)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
2026-07-29 00:07:42 -07:00
Jyothirmai KottuandMick 7c248dde7f [diffusion] fix: don't self-kill diffusion worker when PID 1 is the real parent (#31361)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-07-29 14:35:56 +08:00
siweil 9bdbb180b1 [Disagg][NIXL] Fix heterogeneous attn-TP KV transfer for replicated GQA heads (NIXL_ERR_NOT_FOUND) (#31968) 2026-07-29 14:13:23 +08:00
ef6c07008b Support DCP for Kimi Linear model (#32612)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-07-28 22:59:58 -07:00
Liangsheng Yin c4fc241fd3 [Perf] Free KV pages by segment in the paged allocator without a device sync (#32701) 2026-07-28 22:24:16 -07:00
Xun Sun bf6e80718a [Elastic EP] fix previously flaky test of test_mooncake_ep_small.py (#31706) 2026-07-29 13:14:12 +08:00
Brayden Zhong f01a0c7f97 Fixing MXFP8 online quantization pipeline (#31510) 2026-07-28 21:26:13 -07:00
YC Yen-Ching Tseng 68673fe6c5 [AMD] add Gemma3RMSNorm.forward_hip to unbreak ROCm (#32613) 2026-07-28 21:19:17 -07:00
Mick 9a03bebf13 docs(kimi-k3): clarify VLM compatibility (#32661) 2026-07-29 11:53:04 +08:00
Xiaojun(Robin) Zhang 1af0167493 [EPD][VLM] Fix Kimi-VL 2D encoder grids (#32104)
Signed-off-by: Xiaojun Zhang <zhangxiaojunhust@gmail.com>
2026-07-29 11:46:17 +08:00
Junlin Wu d6fcfe02d6 🐛 [llm][npu][quant] Fix ModelSlim MXFP4 packed weight loading (#32013) 2026-07-29 11:34:41 +08:00
Yihao Wang 9ea964a535 [diffusion] fix: per-shard FP8 scale shape for single-GPU fused linears (#32157) 2026-07-28 20:28:01 -07:00
Baizhou Zhang b21cc8eb9b [misc] Update CodeOwner (#32717) 2026-07-28 19:55:14 -07:00
Liangsheng Yin 14bd315d6e [Refactor] Remove dead allocator backup_state / restore_state (#32709) 2026-07-28 19:49:47 -07:00
monkeyLoveding dac4325c0e sgl-kernel-npu tag update to 2026.7.27 (#32596) 2026-07-29 10:22:13 +08:00
339bef7fad [MLX] Fix overlap-loop request bookkeeping and graceful shutdown (#32447)
Co-authored-by: xiaolin2004 <uwowmhdjwpwpwdhwkw@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-07-28 19:04:53 -07:00
580b1acbe6 [MLX] Move fused swiglu tests to test/registered so CI collects them (#32448)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-07-28 19:04:24 -07:00
Leon GaoandAndrew Gu ee678910f7 [Kernel] Skip KV writes to reserved padding slots (#32477)
Co-authored-by: Andrew Gu <andrew@thinkingmachines.ai>
2026-07-29 09:58:18 +08:00
gjsheu d86492fea0 [NPU] adapt dflash v2 on npu (#31739) 2026-07-29 09:39:19 +08:00
amote-i cb12a1547b [NPU] [DOC] update supported features on ascend npu (#32647) 2026-07-29 09:05:21 +08:00
Lianmin Zheng 16a52bff23 [Refactor] Move sampling tokenizer validation helper (#32694) 2026-07-28 16:48:03 -07:00
YAMYandLee Nau 86ee545388 docs(cookbook): update Kimi-K3 GB200 recipes from measured 4x4 runs (#32592)
Co-authored-by: Lee Nau <lnau@nvidia.com>
2026-07-28 16:26:40 -07:00
Lianmin Zheng 9ca4023b13 [Core] Clean up array-like msgspec structs (#32688) 2026-07-28 16:25:22 -07:00