Commit Graph
15593 Commits
Author SHA1 Message Date
Mohammad Miadh Angkad a2ddf92e61 [CI] Fix Mamba ServerArgs namespace (#32211) 2026-07-23 13:15:50 -07:00
Sam Shleifer 1f9d778d1b Skip dist_init/nccl port prechecks when the dist init method is overridden (#31410) 2026-07-23 12:40:41 -07:00
Rain Jiang 7fe82dd02e create rust workspace (#32014) 2026-07-23 12:02:41 -07:00
Hongkuan Zhou d0b9689805 [Fix] Include disagg prefill waiting queue in FPM (#32122)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-07-23 09:30:27 -07:00
Jinyan Yi b98a577fbe Doc/update ascend quickstart (#32205) 2026-07-23 20:25:20 +08:00
Jialin Ouyangandhzh0425 70ac0c4b0e [UnifiedRadixCache][mamba] Fix mamba state corruption and slot leak when load_back aborts (#30986)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-07-23 18:48:31 +08:00
YAMY c18919f8f3 [Mamba] Add a per-path cap for cached states (#31230) 2026-07-23 17:58:36 +08:00
Baizhou Zhang 20f6a416e7 [Tiny] Skip sm120 deepgemm test temporarily (#32193) 2026-07-23 02:54:58 -07:00
Xiaoyu ZhangandClaude Opus 4.8 62aa85d9aa [Kernel] Sweep missed dedicated kernels into kernels.ops (moe/quant siblings + dspark) (RFC #29630) (#32160)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 17:07:16 +08:00
Liangsheng Yin f35411ee81 [Spec] Enable grammar overlap scheduling for STANDALONE speculative decoding (#32110) 2026-07-23 01:47:27 -07:00
Baizhou Zhang 5387e23ecd [Tiny]Correct runner for testing deepgemm (#32174) 2026-07-23 01:13:14 -07:00
Liangsheng Yin 9b853e6832 [Scheduler] Enable decode retraction ordering under speculative decoding (#32023) 2026-07-23 00:42:56 -07:00
pllimax a25164bda3 Shift nightly image build & pipeline schedule ahead by 2 hours (#32011) 2026-07-23 14:07:53 +08:00
heziiop 235a488c87 [NPU] ascend fuseep use moe ep group (#32040) 2026-07-23 14:07:14 +08:00
silencejade 09071be105 [NPU] [FIX] Fix performance degradation of Qwen3.5-397B-A17B (#32130) 2026-07-23 14:05:10 +08:00
Xiaoyu ZhangandClaude Opus 4.8 11b0e5c5ad [Kernel] Classification cleanup: unify _jit_ naming, drop empty/model groups, add elementwise (RFC #29630) (#32148)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 13:47:02 +08:00
Liangsheng Yin 1b63155efe [Perf] Skip blocks past per-request live length in full-width Triton kernels (#32109) 2026-07-22 22:00:25 -07:00
Yuzhen Zhou eb242b6c03 Seed the GDN CuteDSL correctness test inputs to fix flakiness (#32126) 2026-07-22 21:30:00 -07:00
Xiaoyu ZhangandClaude Opus 4.8 2d1a7be8c4 [Kernel] Reclassify kernel tests by ops group + move helpers out of the package (RFC #29630) (#32128)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 12:18:27 +08:00
jiayisunxandMa Mingfei a2935ce329 [XPU][GDN] add XPU path for causal_conv1d_fn and causal_conv1d_update (#31250)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-23 10:50:20 +08:00
Артем Савкин 108182cb81 [Bugfix] [NPU] Fix w4a8 MoE performance degradation (#32113) 2026-07-23 10:31:32 +08:00
zhaozx-cn 926530b0cf [NPU]remove duplicate code (#31863) 2026-07-23 10:24:31 +08:00
huangtingwei 86e1bb584d [HiCache] Add model-aware key isolation to Mooncake Store (#31920) 2026-07-23 10:20:13 +08:00
Chunyuan WU 60dea26077 [sgl-kernel][CPU] add kernel for shm_allgather_into_tensor and shm_reduce_scatter_tensor (#13397) 2026-07-23 09:19:38 +08:00
YanbingJiang 9a7ac3ecef Fix unnecessary gather/scatter on CPU for non-contiguous Mamba statepool (#31754) 2026-07-23 09:08:45 +08:00
sglang-botandsglang-bot 9bda6fdb9a docs: sync LMSYS SGLang blog cards (#32127)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-23 00:57:01 +00:00
Xiaoyu ZhangandClaude Opus 4.8 99f636a86f [Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 08:35:09 +08:00
Xinyu Jiang 8ce68370b5 [AMD] Build Miles nightly ROCm images with docker/build.py and test ROCm 7.2 (#31757) 2026-07-23 08:30:09 +08:00
danielafrimiandDaniel Afrimi 24da0e51b6 Warn on small Mamba chunked prefill size (#30938)
Co-authored-by: Daniel Afrimi <dafrimi@aws-dfw-cs-001-login-01.cm.cluster>
2026-07-22 15:09:53 -07:00
Mohammad Miadh AngkadandBrayden Zhong 0c29c8fece Bump FlashInfer to 0.6.15.post1 (#31927)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-07-22 14:21:59 -07:00
c20c48b8fd Add 'anyOf' schema support for qwen3_coder tool call parser (#30832)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-07-22 14:14:47 -07:00
Brayden Zhong e7511141ea Support CuteDSL GEMM BF16 on SM100 on by default when allowed by heuristic (#30567) 2026-07-22 14:13:12 -07:00
Liangsheng Yin 98cc8d91cf [Fix] Evict only the KV shortfall in evict_from_tree_cache (#32016) 2026-07-22 12:41:06 -07:00
Liangsheng Yin 5c6a29b3c3 [Fix] Unify pinned host pool release on graceful shutdown (#32029) 2026-07-22 12:29:34 -07:00
Cheng Wan f5dcbe8f14 Revert RuntimeContext config-namespace reads/roles (#31813–#31817) (#32100) 2026-07-22 11:52:41 -07:00
Mohammad Miadh Angkad 0bdd4730af [CI] Fix failures on main (#32091) 2026-07-23 00:11:25 +08:00
Ke Bao 40ac197fb2 Fix get_server_args import lint error (#32096) 2026-07-23 00:09:25 +08:00
Ke Bao e81b326e72 Add Inkling per-commit server test (#32095) 2026-07-22 23:56:18 +08:00
Raghavendra Vedula 4eaa5ca651 Treat partial_json_parser AssertionError as incomplete JSON (#31975) 2026-07-22 23:25:25 +08:00
Raghavendra Vedula a9497e8d73 Guard min_new_tokens penalizer against None eos_token_id (#31973) 2026-07-22 23:24:25 +08:00
Chengze Fan 40b2119b23 [AMD] Cache AITER expert mask across decode (#31889) 2026-07-22 07:50:23 -07:00
e8e765b9d6 [AMD] Add fused all-reduce RMSNorm per-group quant for Qwen3.5 FP8 (#24651)
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
2026-07-22 07:33:03 -07:00
Mohammad Miadh Angkad b855efd9e6 Fix Inkling kernel imports after migration (#32076) 2026-07-22 22:09:01 +08:00
0a6d1930c3 [Attention Backend] Add HPC-Ops attention backend (#30540)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
2026-07-22 22:06:22 +08:00
Fan Kunand范坤 004df6b520 [FIX] Prevent Lightning Attention extra-buffer mamba state corruption (#29973)
Co-authored-by: 范坤 <fankun@U-Q0542J2D-0225.local>
2026-07-22 22:01:29 +08:00
ishandhananiandConnor Carpenter 21065bc862 feat: add native gRPC sidecar module launcher (#31076)
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Signed-off-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
2026-07-22 06:39:08 -07:00
Xiaoyu ZhangandClaude Opus 4.8 74338e94f1 [Kernel] Phase 4 batch-3: migrate tangled JIT subsystems + new groups into kernels.ops (RFC #29630) (#32045)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 21:15:03 +08:00
YC Yen-Ching Tsengandbingxche 71fe649d68 [AMD] Register the Helios release workflow on main (#32069)
Co-authored-by: bingxche <bingxche@amd.com>
2026-07-22 21:14:21 +08:00
b13abfdedf Fix LongCat n-gram token-table crashes on padded batches (#31312)
Co-authored-by: whn09 <whn09@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-07-22 20:24:21 +08:00
Mohammad Miadh Angkad d6690de961 [CI] Fix Marlin MoE test ServerArgs initialization (#32049) 2026-07-22 20:16:43 +08:00