Lianmin Zheng
|
d6ce1b8dce
|
docs: update add-jit-kernel skill for run_suite CI registration (#21264)
|
2026-03-23 21:31:22 -07:00 |
|
Lianmin Zheng
|
260abe1fb1
|
Refactor JIT kernel CI to use run_suite.py registration system (#21239)
|
2026-03-23 21:17:27 -07:00 |
|
  
|
0986bed8e2
|
[HiCache][HybridModel]: Support mamba state offloading & HybridCacheController (#20457)
Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2026-03-23 20:02:50 -07:00 |
|
 Ratish PandXiaoyu Zhang
|
2b1d3c935e
|
[diffusion] fix Z-Image SP sharding for portrait and padded resolutions (#21042)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-03-24 10:15:33 +08:00 |
|
Baizhou Zhang
|
d173cfecd6
|
Update access of cut branch workflow and delete deprecated release workflow (#21231)
|
2026-03-23 15:18:10 -07:00 |
|
Yuxuan Zhang
|
fcaad42b00
|
[Bug Fix] GLM-V / GLM-OCR: field detection for transformers 5.x and MTP omission fix (#21134)
|
2026-03-23 13:19:48 -07:00 |
|
Lianmin Zheng
|
4779755eb9
|
Split pr-test.yml: extract sgl-kernel, jit-kernel, and multimodal-gen tests into separate workflow files (#21219)
|
2026-03-23 13:17:05 -07:00 |
|
Baizhou Zhang
|
ed316a26ef
|
Fix CP in-seq-split method for DeepSeek V32 and update related tests (#21192)
|
2026-03-23 12:34:10 -07:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
27ac831a84
|
docs: improve CI and testing documentation (#21202)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-23 10:48:50 -07:00 |
|
jacky.cheng
|
b4d3fb001d
|
[AMD] Add fused GemmaRMSNorm forward_hip to use aiter/vllm kernels for qwen3.5 (#21188)
|
2026-03-23 10:21:36 -07:00 |
|
Vipul Vaibhaw
|
6d2e1156a0
|
[Test] Add unit tests for srt/constrained module (#21010)
|
2026-03-24 00:33:18 +08:00 |
|
Jiabin Wang
|
83b9d74424
|
ci: unit test for srt/observability module (#21002)
|
2026-03-24 00:30:02 +08:00 |
|
Zijun Gao
|
4dbe42527e
|
[Test] Add unit tests for srt/parser (#20947)
|
2026-03-24 00:26:46 +08:00 |
|
Johnsonms
|
777edb6ef7
|
Fix(jit): support rmsnorm for hidden_size in {64, 128, 256} (#20661)
|
2026-03-23 23:17:44 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
5bdc07d974
|
[Qwen3.5] Fuse split/reshape/cat ops in GDN projection with Triton kernel (#21019)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-03-23 23:17:01 +08:00 |
|
McZyWu
|
8662ba7db4
|
[NPU] bugfix for import sgl-kernel error (#21200)
|
2026-03-23 19:52:36 +08:00 |
|
strgrb
|
80d4a0753a
|
fix fused_set_kv_buffer for rope with Ling-v2 (#20316)
|
2026-03-23 19:20:40 +08:00 |
|
McZyWu
|
4641e5a3d2
|
[NPU] enhance accuracy for model minimaxm2 from 16.5% to 95.5% (#17695)
|
2026-03-23 19:06:38 +08:00 |
|
 XDaoHongandZhengdQin
|
2d288ba8c9
|
[Bugfix] fix npu get kv_item_lens in PD separation when use ASCEND_US… (#15852)
Co-authored-by: ZhengdQin <zhengdqin@gmail.com>
|
2026-03-23 15:56:47 +08:00 |
|
 kpham-sglandClaude Opus 4.6
|
59cb9a9da6
|
[Spec][Ngram] 3/N: Fix synchronization issues in Ngram.cpp (#21186)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-23 00:37:07 -07:00 |
|
Lianmin Zheng
|
814202704b
|
ci: unify PR test suite naming (#21187)
|
2026-03-23 00:18:45 -07:00 |
|
yudian0504
|
3d312643b9
|
[BUGFIX] Fix CP residual size mismatch crash when tp_size == attn_cp_size (#21170)
|
2026-03-23 00:12:58 -07:00 |
|
Lianmin Zheng
|
7757a9ddd0
|
ci: remove IS_BLACKWELL env var; auto-detect Blackwell (#21118)
|
2026-03-22 23:44:48 -07:00 |
|
Even Zhou
|
54fdb357d3
|
[CI][NPU] Fix git report dubious ownership (#21162)
|
2026-03-23 14:36:35 +08:00 |
|
 kpham-sglandClaude Opus 4.6
|
bc4aaab6a1
|
[Spec][Ngram] 2/N: Rename branch length to max trie depth (#21181)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-22 23:35:25 -07:00 |
|
Cheng Wan
|
d6b12c401c
|
Revert "[bugfix] Fix PPMissingLayer AttributeError when Using PP" (#21189)
|
2026-03-22 23:28:36 -07:00 |
|
Zhiqiang Xie
|
13f4f010d8
|
HiSparse for Sparse Attention (#20343)
|
2026-03-22 23:09:31 -07:00 |
|
 Alison ShaoandAlison Shao
|
44db0c59cf
|
[CI] Fix cutlass import error: restore nvidia-cutlass-dsl force-reinstall (#21182)
Co-authored-by: Alison Shao <alison.shao@mac.lan>
|
2026-03-22 22:41:01 -07:00 |
|
Lianmin Zheng
|
7050011dee
|
Enable JIT clamp_position and resolve_future_token_ids on ROCm (#21116)
|
2026-03-22 22:33:54 -07:00 |
|
Yuhao Yang
|
32a85ef128
|
[diffusion] CI: auto-skip diffusion tests when required pipeline class is missing from diffusers (#21139)
|
2026-03-23 12:15:21 +08:00 |
|
Shangming Cai
|
a094eb9c21
|
Temporarily disable flaky qwen3 cp test in CI (#21178)
|
2026-03-22 21:13:52 -07:00 |
|
Mohammad Miadh Angkad
|
d8a5b1dbaf
|
[Bugfix] Work around FlashInfer unified transport issue on GB (#20039)
|
2026-03-22 21:10:25 -07:00 |
|
Xiaoyu Zhang
|
a94d67d44b
|
[SKILL] fix(bench): Support model-specific DenoisingStage variants in… (#21137)
|
2026-03-23 12:08:00 +08:00 |
|
fanghao
|
2b47bd3a34
|
[Bug Fix] Fix non-streaming request abort failure when --enable-metrics is enabled (#20625)
|
2026-03-22 19:58:49 -07:00 |
|
 yuumnandyuumn
|
889e8489e9
|
[diffusion] model: support FireRed-Image-Edit (#20862)
Co-authored-by: yuumn <1010797597@qqã.com>
|
2026-03-23 10:27:07 +08:00 |
|
Cishoon
|
999bad5aba
|
Fix VRAM leak in overlap scheduling with structured output (#20640) (#20697)
|
2026-03-22 17:07:39 -07:00 |
|
Yilong Zhao
|
343998865a
|
perf: pad max-num-requests in decode cuda graph for higher coverage (#20978)
|
2026-03-22 17:06:16 -07:00 |
|
Ziang Li
|
ce0541404f
|
[FlashInfer v0.6.6][RL] Support fp8-last-n-bf16 RL for flashinfer_trtllm_routed moe backend (#20214)
|
2026-03-22 11:17:01 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
c1fe5de69c
|
[Diffusion] Clean up diffusion Triton kernels and modernize custom op registration (#21122)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-22 22:38:57 +08:00 |
|
Ke Bao
|
2406ddfdb8
|
Add ut guide to test skills (#21130)
|
2026-03-22 20:55:37 +08:00 |
|
Brayden Zhong
|
009eee85a0
|
CUTLASS FP8 Blockwise GEMM improvement of SM120 (#20887)
|
2026-03-22 17:55:54 +08:00 |
|
Xiaoyu Zhang
|
766d225fcc
|
Add SGLang CUDA crash API logging inspired by FlashInfer (#20910)
|
2026-03-22 16:39:40 +08:00 |
|
 
|
bb737d7a82
|
Support Qwen3 MoE context parallel (#18233)
Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>
Co-authored-by: Jiying Dong <87510204+dongjiyingdjy@users.noreply.github.com>
|
2026-03-22 01:27:20 -07:00 |
|
kpham-sgl
|
6d160b42bb
|
[Spec][Ngram] 1/N: Reference based Speculative Decoding refactor (#20393)
|
2026-03-22 00:55:10 -07:00 |
|
Liangsheng Yin
|
d9f5c2179c
|
ci(slash-cmd): allow write-permission users to /rerun-ut on fork PRs (#21121)
|
2026-03-22 00:45:48 -07:00 |
|
 Xiaoyu ZhangandMick
|
1b65c0d259
|
[Diffusion] Fix torch.compile RMSNorm fallback for Z-Image (#20962)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-22 15:38:22 +08:00 |
|
Liangsheng Yin
|
1e97864d75
|
ci(slash-cmd): allow repo write-permission users to /rerun-ut (#21120)
|
2026-03-22 00:32:29 -07:00 |
|
 Bowen Liandkinza99
|
3bc595acbc
|
[FlashAttn] Add fused triton kernel for normal_decode_set_metadata (#20778)
Co-authored-by: kinza99 <dh18324568312@163.com>
|
2026-03-22 15:12:29 +08:00 |
|
Mick
|
f7fc2c8592
|
[diffusion] fix: fix accuracy for some image models (#20679)
|
2026-03-22 15:11:57 +08:00 |
|
Liangsheng Yin
|
47d36d73f7
|
Update write-sglang-test skill: CUDA-only for common tests + prefer mock (#21119)
|
2026-03-21 22:54:43 -07:00 |
|