Pavan Sivaram Girijala
|
d9b1cba255
|
Fix DeepSeek-OCR batching crash on variable local-crop counts (#33214)
|
2026-08-05 18:39:19 -07:00 |
|
Zilin Zhu
|
f01f706960
|
[RL] Skip rotary cache tensors in weight checker (#27692)
|
2026-08-05 18:22:12 -07:00 |
|
Eric Zhang
|
ceaeca0b9e
|
[Fix] Vocab out of bounds in DSpark for Inkling-Small (#33748)
|
2026-08-05 18:20:02 -07:00 |
|
 DevashishLal-CBandDevashish Lal
|
fc74c35546
|
kernels: scalar scale A support for fp8_gemm (#33469)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
|
2026-08-06 09:11:25 +08:00 |
|
Eric Zhang
|
269d51ed4b
|
[Fix] Inkling works with gs:// runai_streamer paths (#33750)
|
2026-08-05 18:09:17 -07:00 |
|
 
|
28848bfe7c
|
Fix Nightly NV CI (#33564)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-08-05 18:06:47 -07:00 |
|
 Khoa PhamandClaude Opus 5
|
beabc5949b
|
Enable MoE deferred finalize by default and drop its expert_weights dtype workaround (#33618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-05 17:56:47 -07:00 |
|
 Aditya KamatandBrayden Zhong
|
c9506d023f
|
[Quantization] Route per-tensor FP8 checkpoints to FlashInfer on SM90 (#33148)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-08-06 00:54:10 +00:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4c0a8940fa
|
[Kernel] Unify BaseFusedOp and MultiPlatformOp dispatch (#33205)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 08:52:09 +08:00 |
|
 Xiaoyu ZhangandMohammad Miadh Angkad
|
ba12a16a62
|
[diffusion] Prefer cuDNN SDPA over FA4 for dense attention on sm_100 (B200) (#33655)
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-08-06 08:49:10 +08:00 |
|
 sglang-botandsglang-bot
|
c952ee5ac1
|
chore: bump sgl-kernel version to 0.4.6 (#33678)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-08-05 17:46:54 -07:00 |
|
Mohammad Miadh Angkad
|
65d5a0ec25
|
Support ModelOpt MXFP8 checkpoints (#32538)
|
2026-08-05 17:44:32 -07:00 |
|
Liangsheng Yin
|
3869fe556f
|
[CI] Collapse the EAGLE launch matrix and the scoring engine boots on the per-commit runners (#33756)
|
2026-08-05 16:43:19 -07:00 |
|
Mohammad Miadh Angkad
|
5424d2039c
|
[CI] Fix GLM-Image usage unit tests (#33731)
|
2026-08-05 16:20:08 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
a1cc286062
|
[Intel GPU] DeepSeek V4 4/N: use sgl-kernel implementation of fused_q_norm_rope on XPU (#27790)
Signed-off-by: P V R K Jyothendra Varma <polisettyvarma@gmail.com>
|
2026-08-05 16:04:35 -07:00 |
|
Liangsheng Yin
|
c0ef548eef
|
[misc] Unify MLA scaling init and remove dead buffer / scaling code (#33363)
|
2026-08-05 15:59:33 -07:00 |
|
YAMY
|
990a446773
|
Fix padded positions in breakable CUDA Graph attention (#33253)
TestBreakableCUDAGraphBasic and all NVIDIA CI tests pass.
|
2026-08-05 15:53:28 -07:00 |
|
 YAMYandChunan Zeng
|
7bc90ab394
|
Select DeepGEMM standard layouts by memory budget (#33474)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
|
2026-08-05 15:50:43 -07:00 |
|
 yvbbrjdrandhnyls2002
|
02cd44c59a
|
feat(dgx-spark): add inkling-small MoE support for sm_121 (#33108)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-08-05 15:06:59 -07:00 |
|
cctry
|
a3a1ebc7b7
|
Warn on risky serving-time Triton work (#33120)
|
2026-08-05 14:55:03 -07:00 |
|
Ziang Li
|
988c6e6aeb
|
Pin online NVFP4 4over6 quantization settings (#33621)
|
2026-08-05 21:29:32 +00:00 |
|
 Brayden ZhongandBrayden Zhong
|
a14c870886
|
Fix broken Nemotron DP attention (#33123)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-08-05 14:27:05 -07:00 |
|
Hank Han
|
9436de717f
|
[Spec][PD] Enable fused TopK for GLM-5.2 MTP IndexShare (#31477)
|
2026-08-05 14:17:42 -07:00 |
|
Zhiqiang Xie
|
106bcc1293
|
Observability enhancement for HiCache (#32388)
|
2026-08-05 14:13:06 -07:00 |
|
Shu Wang
|
55b1c09e73
|
[core] Consolidate compiled-kernel caches under SGLANG_CACHE_DIR (#32434)
|
2026-08-05 13:54:27 -07:00 |
|
Jialin Ouyang
|
717a559f02
|
[Scheduler] Align WAR fences with CUDA graph metadata reads (#33587)
|
2026-08-05 13:52:20 -07:00 |
|
 ![copilot-swe-agent[bot]](/assets/img/avatar_default.png)   
|
ea65f8ddc9
|
Feat/spectrum (#31491)
Co-authored-by: jaron1990 <jaron1990@gmail.com>
Co-authored-by: jaron1990 <34618972+jaron1990@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: alexnails <alex.nails@radixark.ai>
|
2026-08-05 13:23:17 -07:00 |
|
 
|
5c4f72f92a
|
[Build] Add srt_empty extra group for device-agnostic install (#31300)
Co-authored-by: yixiaodapeng <yixiaodapeng@example.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-05 13:17:46 -07:00 |
|
Liangsheng Yin
|
de34dd11e9
|
[CI] Fold duplicate-server suites and prune the retract matrix on 1-gpu-5090 (#33745)
|
2026-08-05 12:41:51 -07:00 |
|
Jason Mancuso
|
36853b8ffc
|
[Spec] Support logprobs with DFlash (#33459)
|
2026-08-05 12:37:12 -07:00 |
|
Khoa Pham
|
5f79cf3511
|
[DCP] Match the replicated draft KV pool's page granularity to its allocator (#33348)
|
2026-08-05 11:40:28 -07:00 |
|
Jialin Ouyang
|
b1bd871df5
|
[Unified Radix Cache] Complete the tree-core interface boundary (#33580)
|
2026-08-05 11:39:28 -07:00 |
|
cctry
|
96c89863a3
|
Measure prefill busy time between launches (#33595)
|
2026-08-05 11:26:27 -07:00 |
|
WenhaoZhang
|
acaab22d09
|
[diffusion] feat: add SageAttention packed varlen path for minimax-h3 (#33703)
|
2026-08-06 01:19:58 +08:00 |
|
zhaozx-cn
|
4e7209caa8
|
[NPU] Add causal conv1d (#28267)
|
2026-08-05 22:22:49 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
3425c93666
|
[diffusion] Wan VAE RMSNorm+SiLU fusion behind quality=high (H200 FastWan2.2 e2e 9.611 -> 9.125 s) (#33546)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-05 21:33:35 +08:00 |
|
silencejade
|
593777c046
|
[FIX] [benchmark] Fix flush_cache failure after warmup by waiting for server idle (#33527)
|
2026-08-05 21:27:43 +08:00 |
|
 Xuan LiaoandMa Mingfei
|
3b4fac5b99
|
[XPU] DeepSeek V4: use sgl-kernel-xpu implemetation of flash_mla_sparse_fwd for prefill (#31865)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-08-05 21:05:53 +08:00 |
|
Mick
|
99709f734d
|
[VLM] split multimodal scheduling from mm_utils (#32415)
|
2026-08-05 20:24:12 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
a5888c956f
|
[diffusion] Pack Ulysses Q/K/V input all-to-all into one collective + reusable a2a staging buffers (#33667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-05 19:15:46 +08:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
2f22ed58ea
|
[NPU] Adding a fast layernorm for diffusion models and fix BSA (#29027)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-08-05 14:06:00 +03:00 |
|
 
|
22d558b103
|
[Feature] Add GLM Image usage report (#33378)
Co-authored-by: wuyuefeng <wuyuefeng@noreply.gitcode.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-08-05 13:59:49 +03:00 |
|
 Alex NailsandClaude Fable 5
|
6fa3f9df11
|
[Bugfix] Treat unsharded model.safetensors as HF weights in Mistral-native format detection (#33671)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-05 01:54:46 -07:00 |
|
 
|
a6e5fa7081
|
[Scheduler] Honor explicit min-free-slots thresholds (#33403)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-08-05 01:44:18 -07:00 |
|
Trevor Morris
|
81c7a54ecd
|
[NVIDIA] Use sm_100f instead of sm_100a for sgl-kernel and FlashMLA (#33433)
|
2026-08-05 01:36:46 -07:00 |
|
Xinyi Song
|
1478cdec9f
|
[AMD] Fuse Kimi-K3 attn-residual aggregation (#33599)
HIP Gated changes
|
2026-08-04 23:20:04 -07:00 |
|
 Артем СавкинandXiaoyu Zhang
|
d96df7bed5
|
[Diffusion] Batch GLM-Image AR requests (#30683)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-05 08:47:06 +03:00 |
|
 
|
059269594c
|
[DSV4] Add official DSV4 reasoning effort support (#33140)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: David Orman <ormandj@corenode.com>
|
2026-08-05 12:50:41 +08:00 |
|
Liangsheng Yin
|
198a3bc29b
|
[Test] Route GEMM backend UTs through real layer modules and weight loaders (#33615)
|
2026-08-04 20:53:26 -07:00 |
|
Liangsheng Yin
|
1033cae8d5
|
[CI] Speed up dependency install: dual-ABI Rust ext cache and prevalidation pruning (#33619)
|
2026-08-04 20:33:48 -07:00 |
|