Ke Bao
|
3974b00359
|
Add bit-exact unified radix cache KL test for hybrid SWA + mamba (#34607)
|
2026-08-13 02:15:36 +08:00 |
|
   
|
773faf992d
|
Reserve multimodal runtime allocations and keep padded inputs aligned (#34141)
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: wangwenchen0407 <wangwenchen@meta.com>
Co-authored-by: Hanming Lu <hanminglu@meta.com>
|
2026-08-12 11:04:07 -07:00 |
|
 
|
b501311fa1
|
[Kimi K3] Fuse MLA gate projection into QKV-A GEMM (#33623)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-08-13 00:21:34 +08:00 |
|
Mick
|
ad47dde65c
|
[diffusion] optimize: fuse cosmos qk norm, rope, and kv packing (#34275)
|
2026-08-12 23:18:16 +08:00 |
|
Mohammad Miadh Angkad
|
00e57d74f0
|
Bump FlashInfer to 0.6.17 and remove Kimi K3 workarounds (#33997)
|
2026-08-12 02:17:26 -07:00 |
|
Yuzhen Zhou
|
2d76d537e5
|
feat: support deterministic FA4 for GLM-4.7-Flash (#33945)
|
2026-08-12 16:57:34 +08:00 |
|
Xiaoyu Zhang
|
1f008dc226
|
[Diffusion][LTX-2] Allocate AdaLN outputs from one contiguous slab (#34508)
|
2026-08-12 16:28:27 +08:00 |
|
 Sam ShleiferandClaude Fable 5
|
a2e88279c2
|
metrics: don't clock-rebase unset time sentinels in ReqTimeStats deserialization (#34335)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-12 14:27:47 +08:00 |
|
+3        
|
5899674504
|
[Fix] Make DeepSeek-V4 reasoning and tool-call streaming parsing chunk-invariant (#34458)
Co-authored-by: hao-cyber <89575785+hao-cyber@users.noreply.github.com>
Co-authored-by: Enrico Falco <enrico9034@gmail.com>
Co-authored-by: Svyatoslav <85786374+slivanovich@users.noreply.github.com>
Co-authored-by: Andreas Hassellof <andreas@ombori.com>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: Chenglun Hu <chenglunhu@gmail.com>
Co-authored-by: robellliu-dev <robell.liu@huawei.com>
Co-authored-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: tancheng33 <garrytancheng@gmail.com>
Co-authored-by: dineshx29 <dinesh.b.offl@gmail.com>
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
|
2026-08-11 20:03:28 -07:00 |
|
Xiaoyu Zhang
|
4aff4b1822
|
[Diffusion] Improve bit-exact fusion fallback diagnostics (#34412)
|
2026-08-12 10:42:48 +08:00 |
|
zijiec
|
93e9db5eb8
|
[Fix][Qwen]: fused shared-expert detection PP-safe protection (#34447)
|
2026-08-11 18:50:05 -07:00 |
|
Yanbin Jiang
|
1ce515a53d
|
Refocus LoRA tests on regression coverage (#34464)
|
2026-08-11 17:37:52 -07:00 |
|
Shu Wang
|
c7c03ec53b
|
[NVIDIA] Add flashinfer MNNVL backend for allreduce only (#30700)
|
2026-08-11 16:46:46 -07:00 |
|
 
|
d82a1d4802
|
[XPU] Pad MoE expert weight row stride to avoid L3 aliasing (#33905)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-11 15:54:13 -07:00 |
|
     
|
fde9ad2531
|
[Feature] Add Muse Glimmer model support (#34262)
Co-authored-by: sglang-bot <232288953+sglang-bot@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-08-11 15:41:52 -07:00 |
|
Mohammad Miadh Angkad
|
7fb6e61b95
|
Fix CUDA 13.0 VMM handle type compatibility (#34431)
|
2026-08-11 15:16:34 -07:00 |
|
ilyasher-harmonic
|
9ced8d0981
|
Optimize FP32 LM head for bf16/fp16 (#32370)
|
2026-08-11 15:14:50 -07:00 |
|
gongwei1027
|
2c07ca5e8d
|
[Fix] Allow flashinfer_sparse_mla DSA backend for HiSparse on SM120 FP8 KV (#33075)
|
2026-08-11 15:05:04 -07:00 |
|
 Dmitrii SergeevandZhiqiang Xie
|
c58953d90a
|
O(1) slot allocation in ReqToTokenPool.alloc() (#32208)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2026-08-11 11:26:05 -07:00 |
|
Jeremy Zhang
|
aadb9720fe
|
fix: route scheduler aborts to multi-tokenizer workers (#33940)
|
2026-08-11 11:15:16 -07:00 |
|
Ke Bao
|
f9153df62e
|
Add bit-exact hicache logprob-consistency test (#34356)
|
2026-08-12 01:54:45 +08:00 |
|
 huangtingweiandHanming Lu
|
8f3d3a31f4
|
[HiCache] Fix Mamba track-boundary bookkeeping under overlap scheduling (#29792)
Co-authored-by: Hanming Lu <hanminglu@meta.com>
|
2026-08-12 00:36:48 +08:00 |
|
Mick
|
8267d76c2c
|
[VLM] replace deprecated image processor use_fast (#34175)
|
2026-08-12 00:14:07 +08:00 |
|
Ke Bao
|
b20c375c10
|
Fix flaky decode cache-hit check in Inkling test (#34405)
|
2026-08-11 21:47:01 +08:00 |
|
Mohammad Miadh Angkad
|
2d193077f7
|
[JIT Kernel] Migrate per-token FP8 quantization from AOT to JIT (#34257)
|
2026-08-11 20:40:40 +08:00 |
|
 Zhiqiang XieandTingwei Huang
|
5469faec45
|
HiSparse: shared-index (IndexShare) plan-then-IO swap-in prefetch (#34329)
Co-authored-by: Tingwei Huang <huangtingwei9988@gmail.com>
|
2026-08-11 01:58:28 -07:00 |
|
 
|
e74ea5b1d7
|
[ROCm/gfx95] Fix fp8 per-channel attention for Kimi-K2.7-code-mxfp4 o… (#31105)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-08-11 00:51:57 -07:00 |
|
ybyang
|
9d4be40124
|
Fix DSpark + DeepSeek V4 prefill CP compatibility (#33865)
|
2026-08-10 23:26:29 -07:00 |
|
 YAMYandShangming Cai
|
667e18d99d
|
[PD] Support pipeline-parallel prefill with Mooncake staging buffer (#33807)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-08-11 13:53:17 +08:00 |
|
EchO
|
a58fa0388e
|
[Fix] Correct W4AFP8 DeepEP scaling and mode-specific dtypes (#33669)
|
2026-08-11 11:21:42 +08:00 |
|
Jae B.
|
5af5183351
|
test: isolate metal profiler tests from ambient SGLANG_USE_MLX (#34300)
|
2026-08-10 18:44:16 -07:00 |
|
cctry
|
df986c4d5e
|
Consolidate CUDA VMM allocation helpers (#34199)
|
2026-08-10 18:11:11 -07:00 |
|
Mick
|
418975ba64
|
[EPD] feat: pipeline owner-only multimodal preprocessing (#34206)
|
2026-08-11 09:08:05 +08:00 |
|
milesial
|
d59c1ddf70
|
fix(dflash): account for DCP in draft KV pool sizing (#33912)
Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com>
|
2026-08-10 17:15:01 -07:00 |
|
 Khoa PhamandClaude Opus 5
|
0967885121
|
[DCP] Drop two per-layer launches from the MLA target-verify path (#34240)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-10 17:11:03 -07:00 |
|
Mohammad Miadh Angkad
|
8c5d5f75bf
|
Fix DSV4 DSpark shared expert loading (#33312)
|
2026-08-11 07:36:50 +08:00 |
|
Mohammad Miadh Angkad
|
77b8315b84
|
[CI] Fix GSM8K floating-point tolerance boundary (#34272)
|
2026-08-10 16:27:11 -07:00 |
|
 Cheng WanandClaude Fable 5
|
ceeaec2078
|
[PD] Support --enable-unified-memory with PD disaggregation (kimi-linear MLA hybrid-Mamba) (#33362)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-10 16:07:58 -07:00 |
|
 Cheng WanandClaude Fable 5
|
8a7c8a72d6
|
Fix NaN logits from deterministic Triton extend on the unified memory pool (#33517)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-10 16:05:04 -07:00 |
|
AMD-yanfeiwang
|
ca0f8a0f4c
|
perf(hisparse): fuse the DSv4 value and scale swap-in copy on ROCm (#33484)
|
2026-08-10 14:30:15 -07:00 |
|
AMD-yanfeiwang
|
1a8e4876b6
|
perf(hisparse): 128-bit non-temporal swap-in copy on ROCm (#33085)
|
2026-08-10 13:27:15 -07:00 |
|
 Lifan ShenandXinyuan Tong
|
a2161ce682
|
Support thinking budget for Inkling (#33146)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-08-11 02:01:08 +08:00 |
|
Cheng Wan
|
7738062294
|
[unified memory] Support DSPARK speculative decoding + fix two NaN root causes (page hand-out zeroing, CuTe int32 slot-stride wrap) (#33974)
|
2026-08-10 10:35:06 -07:00 |
|
 
|
fd3036523a
|
[diffusion] Clean up shared bitexact gates, helpers, and stale naming (#34180)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-10 22:14:09 +08:00 |
|
 
|
14ffd447a4
|
[FEAT] Decouple multimodal global cache from Mooncake (#30392)
Co-authored-by: Yuang Chen <cya539102@antgroup.com>
Co-authored-by: Yuang Chen <1131578721@qq.com>
|
2026-08-10 19:23:19 +08:00 |
|
Mick
|
443b62db57
|
fix(vlm): stream-order cuda-ipc feature pool lifecycle and streamline multimodal transport module (#33949)
|
2026-08-10 18:47:19 +08:00 |
|
 YAMYandShangming Cai
|
c971d7ac9c
|
Refactor staging registration metadata fields (#33910)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-08-10 17:53:32 +08:00 |
|
Jincong Chen
|
3c533acec6
|
[Hicache][2/2]Support Mamba branching in Unified Radix Cache with HiCache (#33639)
|
2026-08-10 17:16:42 +08:00 |
|
 
|
a76a167812
|
Fix/hisparse host backed max request length (#28753)
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-08-10 16:58:13 +08:00 |
|
 Liangsheng YinandBrayden Zhong
|
b51bf9ec9e
|
[Spec] Budget the DFLASH draft KV pool from its own attention geometry (#34234)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-08-10 01:28:49 -07:00 |
|