18694 Commits
Author SHA1 Message Date
Xinyuan TongandZijie Xia b3cdd016ba Add Ling-3.0-flash cookbook (#33556)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-05 22:53:34 +08:00
zhaozx-cn 4e7209caa8 [NPU] Add causal conv1d (#28267) 2026-08-05 22:22:49 +08:00
Xiaoyu ZhangandClaude Fable 5 3425c93666 [diffusion] Wan VAE RMSNorm+SiLU fusion behind quality=high (H200 FastWan2.2 e2e 9.611 -> 9.125 s) (#33546)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 21:33:35 +08:00
silencejade 593777c046 [FIX] [benchmark] Fix flush_cache failure after warmup by waiting for server idle (#33527) 2026-08-05 21:27:43 +08:00
Xuan LiaoandMa Mingfei 3b4fac5b99 [XPU] DeepSeek V4: use sgl-kernel-xpu implemetation of flash_mla_sparse_fwd for prefill (#31865)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-05 21:05:53 +08:00
Mick 99709f734d [VLM] split multimodal scheduling from mm_utils (#32415) 2026-08-05 20:24:12 +08:00
Xiaoyu ZhangandClaude Fable 5 a5888c956f [diffusion] Pack Ulysses Q/K/V input all-to-all into one collective + reusable a2a staging buffers (#33667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 19:15:46 +08:00
2f22ed58ea [NPU] Adding a fast layernorm for diffusion models and fix BSA (#29027)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-05 14:06:00 +03:00
22d558b103 [Feature] Add GLM Image usage report (#33378)
Co-authored-by: wuyuefeng <wuyuefeng@noreply.gitcode.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-05 13:59:49 +03:00
YC Yen-Ching TsengandChen 8279702e0b [AMD] Stop publishing the K3 MI35X nightly image (#33689)
Co-authored-by: Chen <bingxche@amd.com>
2026-08-05 17:53:23 +08:00
Alex NailsandClaude Fable 5 6fa3f9df11 [Bugfix] Treat unsharded model.safetensors as HF weights in Mistral-native format detection (#33671)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 01:54:46 -07:00
Liangsheng Yin 4a3d6ca88c [CI] Skip sglang-kernel and sgl-deep-gemm reinstall on version match (#33637) 2026-08-05 01:46:16 -07:00
a6e5fa7081 [Scheduler] Honor explicit min-free-slots thresholds (#33403)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-08-05 01:44:18 -07:00
Liangsheng Yin c0d5ebd6c4 [CI] Move CPU-only unit tests to the CPU suite and trim dead 5090 registrations (#33654) 2026-08-05 01:43:29 -07:00
Mohammad Miadh Angkad 98ed5554bb Stop testing cu129 DeepGEMM wheels (#33675) 2026-08-05 01:38:28 -07:00
Trevor Morris 81c7a54ecd [NVIDIA] Use sm_100f instead of sm_100a for sgl-kernel and FlashMLA (#33433) 2026-08-05 01:36:46 -07:00
Xinyi Song 1478cdec9f [AMD] Fuse Kimi-K3 attn-residual aggregation (#33599)
HIP Gated changes
2026-08-04 23:20:04 -07:00
Артем СавкинandXiaoyu Zhang d96df7bed5 [Diffusion] Batch GLM-Image AR requests (#30683)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-08-05 08:47:06 +03:00
Liangsheng Yin eac1f78568 [CI] Free hosted-runner disk space only when it is low (#33644) 2026-08-04 21:55:06 -07:00
059269594c [DSV4] Add official DSV4 reasoning effort support (#33140)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: David Orman <ormandj@corenode.com>
2026-08-05 12:50:41 +08:00
Liangsheng Yin 198a3bc29b [Test] Route GEMM backend UTs through real layer modules and weight loaders (#33615) 2026-08-04 20:53:26 -07:00
Liangsheng Yin 1033cae8d5 [CI] Speed up dependency install: dual-ABI Rust ext cache and prevalidation pruning (#33619) 2026-08-04 20:33:48 -07:00
Kaixi 6c05aaae7e [trtllm_mha] perf: Stop allocating per-layer scratch inside the decode CUDA graph (#33063) 2026-08-04 19:27:26 -07:00
4949b5fccf [XPU] Add qknorm_rope support for Flux (#30883)
Co-authored-by: Chandrakant Khandelwal <Chandrakant.Khandelwal@intel.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-05 10:12:24 +08:00
gjsheu 5dc4102960 [npu] [bugfix] Fix PD‑disaggregation error (#33523) 2026-08-05 09:54:00 +08:00
a0b04dbe4c feat(grpc): add generation request semantics (#32588)
Signed-off-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-04 18:53:45 -07:00
Zaili WangandMa Mingfei 29831d58ef fix mm-chunk-embedding test suite (#32895)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-05 09:49:39 +08:00
Polisetty V R K Jyothendra VarmaandMa Mingfei d2c405f19d [Intel GPU] DeepSeek V4 8/N: use sgl-kernel implementation of fused_k_norm_rope_flashmla on XPU (#28040)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-05 09:28:43 +08:00
Qiaolin Yu 9303e26f03 [ci] add qwen 3.5 mtp + replayssm + flashinfer gdn test (#33607) 2026-08-04 18:08:42 -07:00
211ee64249 [rotary] Rebuild the shared RoPE cache entry when its buffers are dead (#33575)
Co-authored-by: mxz <mxz@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-08-04 17:24:42 -07:00
Baizhou Zhang 87ed82ff7e Remove custom all-reduce disable from Kimi-K3 B300 recipe (#33612) 2026-08-04 16:07:31 -07:00
Liangsheng Yin 76dc89f5aa [Test] Replace NVFP4 MoE runner backend e2e matrix with a layer-level unit test (#33611) 2026-08-04 16:03:29 -07:00
Liangsheng Yin 0d99d91e49 [CI] Make B200 base-b suites single-GPU as prep for 1-gpu B200 runners (#33605) 2026-08-04 15:53:01 -07:00
Liangsheng Yin a0b3f1dde6 [Test] Replace GEMM backend e2e matrices with layer-level unit tests (#33596) 2026-08-04 15:50:41 -07:00
Lianmin Zheng b0fd31ba07 Multiple flexibility fixes for DP attention (#33537) 2026-08-04 15:40:40 -07:00
Baizhou Zhang 6808c6d571 [Tiny] Little enhancement of Kimi-K3 test (#33609) 2026-08-04 15:36:41 -07:00
Liangsheng Yin 58da9859c4 [CI] Extract download-rust-ext and give every install step a cache fallback (#33597) 2026-08-04 15:33:34 -07:00
Zhaoyi Li b327d76682 [AMD] Bump mori to latest in sglang (#33462) 2026-08-04 15:12:40 -07:00
cctry c8822fd990 Clarify post-capture KV reservation logs (#33598) 2026-08-04 15:06:11 -07:00
FilipandClaude Opus 4.8 19d3f86895 [LoRA] Laguna: per-layer LoRA hidden-dim resolution for packed attention (#30298)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-04 14:48:05 -07:00
Xinyuan TongandAlex Nails a9c3b55435 [Refactor] Keep chat template validation out of ServerArgs dispatcher (#33392)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-04 14:45:47 -07:00
Jimmy ShongandClaude Opus 5 e510dc58ba Add @Jiminator as codeowner for Laguna model and config (#33472)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 14:10:19 -07:00
Lianmin Zhengandcctry 34af3ff386 Allow optimistic prefill with L2 hierarchical cache and write-back policy (#33545)
Co-authored-by: cctry <cctry@meta.com>
2026-08-04 13:23:08 -07:00
+26 abddb1c7e9 [Kimi] Support kimi-k3 (#32541)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai>
Co-authored-by: Zijie Xia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: zhangxiaohao <1024393531@qq.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: RolaoDenthu <xinyisong0111@gmail.com>
Co-authored-by: pigeonsoup <32922982+pigeonsoup@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com>
Co-authored-by: Lee Nau <lee.nau@gmail.com>
Co-authored-by: HMING <126185151+Hearum@users.noreply.github.com>
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: Hanming Lu <hanminglu@meta.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
2026-08-04 13:22:49 -07:00
Liangsheng Yin 0753663b8e [CI] Trim redundant B200 test registrations (#33586) 2026-08-04 13:22:00 -07:00
Xingyu Liu aa06433709 Avoid TRTLLM prefill output copy (#33306) 2026-08-04 12:54:04 -07:00
38dc2d6cf8 [metrics] Split tokenizer request metrics by stream (#32734)
Co-authored-by: wpc <wpc@devvm23443.cco0.facebook.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-08-04 12:53:52 -07:00
Lianmin Zheng 4794b401d5 [Observability] Add startup, memory, and hybrid SWA diagnostics (#33375) 2026-08-04 12:50:09 -07:00
Lianmin ZhengandJialin Ouyang 5081c063c0 fix(metrics): clear forward occupancy on idle (#33562)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-08-04 12:49:17 -07:00
Lianmin ZhengandItai Gat dea2be5ae3 [CUDA Graph] Allow custom decode graph runners (#33553)
Co-authored-by: Itai Gat <itaigat.mail@gmail.com>
2026-08-04 12:48:56 -07:00