 Thomas WangandBingxu Chen
|
f5b041622b
|
[AMD] Fix deepseek-v4 mtp accept length issue (#28520)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-06-17 11:13:55 -07:00 |
|
 kousakawangandkousakawang
|
8aaca72c21
|
[FIX]Fix Step3-VL multi-image embedding and local patch splitting (#24970)
Co-authored-by: kousakawang <wanghanpei@bytedance.com>
|
2026-06-17 10:31:32 -07:00 |
|
Junlin Wu
|
873196f7fa
|
♻️ [llm][npu][quant] Delegate MXFP8 dense scheme to kernel and use torch.ops.npu (#28505)
|
2026-06-17 10:18:26 -07:00 |
|
Mick
|
735a256f98
|
[diffusion] feat: use LocalAttention for mistral3 encoder (#28176)
|
2026-06-17 21:18:41 +08:00 |
|
 Aleksi VesantoandMick
|
dad890fff1
|
[diffusion] perf: shard text when using sp in flux.1/2 (#27066)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-17 21:17:39 +08:00 |
|
Ziang Li
|
3fb65ebabd
|
[RL] Fix FlashInfer TRTLLM MXFP8 dense weight layout (#28459)
|
2026-06-17 10:33:35 +00:00 |
|
lmyybh
|
2f1390fcb1
|
fix: preserve divisible FP8 block K configs on CUDA (#27553)
|
2026-06-17 10:30:56 +00:00 |
|
YC Yen-Ching Tseng
|
9b8c41171a
|
[AMD] Fall back to layer_first layout for kernel write-back on ROCm (#28473)
|
2026-06-17 03:01:43 -07:00 |
|
Thomas Wang
|
21a95333d4
|
[AMD] Add transpose_scale arg for o_proj to fix GLM accuracy issue (#27798)
|
2026-06-17 01:01:02 -07:00 |
|
Liangsheng Yin
|
f86e9b48e8
|
[Perf] Make spec-decode penalty H2D non-blocking and share decode cumulate path (#28500)
|
2026-06-17 00:59:58 -07:00 |
|
 Ryan Zzzandzhujunyu
|
8fd1694dd2
|
Deepseek v4: support mixed dtype compression states (#27277)
Co-authored-by: zhujunyu <zhujunyu.666@bytedance.com>
|
2026-06-17 00:56:41 -07:00 |
|
 
|
7256ee9871
|
[AMD] Update test_aiter_allgather_amd.py data types alignment between benchmark aiter and custom all-reduce kernel (#27815)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2026-06-17 00:48:15 -07:00 |
|
Peng Xingchen
|
d0e974f40b
|
[NPU] Use use_dsa to dispatch Ascend DSA attention (#28436)
|
2026-06-17 15:47:12 +08:00 |
|
 
|
b54f8432ad
|
Batch EAGLE draft/draft-extend replay memcpys via grouped foreach copy (#28465)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
|
2026-06-17 00:34:21 -07:00 |
|
Liangsheng Yin
|
3bc618485a
|
[Perf] Make latest_output_ids H2D non-blocking in prepare_for_decode (#28491)
|
2026-06-16 23:46:35 -07:00 |
|
Baizhou Zhang
|
27291118b9
|
Upgrade sgl-deep-gemm to 0.1.3 (#28402)
|
2026-06-16 23:06:36 -07:00 |
|
Oxana Korzh
|
c01f62e341
|
[bugfix] guard NVIDIA SM-capability checks with is_cuda() for AMD/ROCm (#28486)
|
2026-06-16 22:51:42 -07:00 |
|
Mick
|
0f5e14e1d9
|
[diffusion] fix: use Megatron-style tp for native encoders and dits (#28318)
|
2026-06-17 13:07:44 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Rohit Kumar Singhandgithub-actions[bot]
|
9371062ef3
|
Fix deep seek ocr2 image processing (#27884)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-06-17 12:59:18 +08:00 |
|
Michael
|
a2aa51c818
|
[AMD] register 4 2-gpu tests to stage-b-test-2-gpu-large-amd (#28344)
|
2026-06-16 21:58:34 -07:00 |
|
Chengze Fan
|
66ac385f52
|
fix(moe): MoRI EP init_mori_op missing BF16 dispatch branch (#28469)
Signed-off-by: Chengze Fan <fancz2002@gmail.com>
|
2026-06-16 21:02:06 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
b8a73bfba0
|
Call Flashinfer mm_fp8 for per-tensor FP8 GEMMs on SM100 (#28333)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-16 20:50:03 -07:00 |
|
 Xinyuan TongandZijie Xia
|
72ccfec594
|
docs(cookbook): verify GLM-5.2 single-node B300 (FP8 + BF16) (#28460)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-06-17 03:47:32 +00:00 |
|
Mick
|
c17190c059
|
update codeowners (#28478)
|
2026-06-17 11:16:52 +08:00 |
|
 Yufeng HeandYufeng He
|
6c8fdb5b62
|
[diffusion] fix: fix PicklingError with --backend diffusers on non-T2I models (#21472)
Co-authored-by: Yufeng He <40085740+universeplayer@users.noreply.github.com>
|
2026-06-17 11:14:25 +08:00 |
|
 Kangyan-ZhouandClaude Opus 4.8
|
827fc56e00
|
[router] Multi-arch experimental sgl-router image + fix distroless libpcre2 startup crash (#28474)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-16 20:09:39 -07:00 |
|
Zilin Zhu
|
f06e2d3d1f
|
Support asymmetric compressed-tensors MoE (#27690)
|
2026-06-16 19:53:33 -07:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) zhaozx-cnandgithub-actions[bot]
|
224b1dc775
|
[NPU]Replace ascend vision attn operator (#25768)
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-06-17 10:33:37 +08:00 |
|
Qeeweew
|
0ae4740bd1
|
fix: add missing clamp_limit for CompressedTensorsWNA16MoE (#27328)
|
2026-06-16 19:29:57 -07:00 |
|
 Trevor MorrisandClaude Opus 4.7
|
9c53853ea3
|
Use pack topk ids triton kernel for flashinfer_trtllm_routed (#25702)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-16 19:27:14 -07:00 |
|
Yanbin Jiang
|
093908d4c0
|
[LoRA] Fix chunked SGMV (csgmv) CUDA graph segment replay (#28371)
|
2026-06-16 19:19:07 -07:00 |
|
chenxu214
|
71b090a8e7
|
[Ascend]GLM 5.2 deployment (#28433)
|
2026-06-17 09:55:50 +08:00 |
|
amote-i
|
74155d32bd
|
[NPU] [DOC] Update Ascend NPU docs: HDK 25.5.2, Triton 3.2.1.dev20260530 (#28311)
|
2026-06-17 09:38:31 +08:00 |
|
Thomas Wang
|
0d651e653b
|
[AMD] Update v4 amd cookbook (#28423)
|
2026-06-16 18:15:16 -07:00 |
|
 
|
37ef295c78
|
[AMD] Feat/dp moe reduce scatter (#28216)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Wang, FangYuan <39615225+At1a8@users.noreply.github.com>
|
2026-06-16 18:12:39 -07:00 |
|
jasonjk-park
|
d86a7e7018
|
Custom spec algorithm can handle server args (#28162)
|
2026-06-16 17:13:52 -07:00 |
|
Jia-Wei Jiang
|
ca84d52b78
|
[Router] [Docs] Refresh policy-selection notes (#28321)
Signed-off-by: JiangJiaWei1103 <waynechuang97@gmail.com>
|
2026-06-16 17:07:19 -07:00 |
|
Qiaolin Yu
|
2ad00faae1
|
[ci] add kimi nvfp4 nightly tests (#28467)
|
2026-06-16 16:49:51 -07:00 |
|
ashwini rathi
|
4f9b12c5dd
|
[XPU] Guard tvm_ffi import in dsv4 compress modules under TYPE_CHECKING (#28426)
|
2026-06-16 15:52:12 -07:00 |
|
Hanming Lu
|
a10eee3d80
|
[Tokenizer] Fix abort racing server crash when large amount of aborts (#28341)
|
2026-06-16 14:49:35 -07:00 |
|
Jyothirmai Kottu
|
b8b8992dde
|
docs: add Amazon SageMaker AI deployment guide (#28338)
|
2026-06-16 13:33:22 -07:00 |
|
huangtingwei
|
9b4432fe18
|
[HiCache]Asymmetric pool support direct backend (#28446)
|
2026-06-16 13:17:57 -07:00 |
|
Jason Mancuso
|
c0a6c3ce66
|
Fix circular import when sglang.srt.model_executor.runner_backend is imported first (#28002)
|
2026-06-16 12:04:34 -07:00 |
|
 
|
13537f8e20
|
Unskip Marlin NVFP4 tests (#27589)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: shaunkotek <shaunkotek@users.noreply.github.com>
|
2026-06-16 11:58:22 -07:00 |
|
Xinyuan Tong
|
33f205d8c5
|
docs(cookbook): fix GLM-5.2 thinking toggle kwarg + document reasoning effort (#28454)
|
2026-06-16 18:17:34 +00:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
799584e173
|
fix: get_processor fails when --tokenizer-path lacks model config.json (#25643)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-16 20:25:05 +03:00 |
|
 feliang-gitandxutizhou
|
92b42c8d8a
|
LPLB: linear-programming load balancer for MoE expert parallelism (#24515)
Co-authored-by: xutizhou <xutingz@nvidia.com>
|
2026-06-16 10:19:42 -07:00 |
|
Xinyuan Tong
|
00081a00d5
|
docs(cookbook): tune GLM-5.2 MTP to 5-1-6 and simplify launch flags (#28448)
|
2026-06-17 01:18:34 +08:00 |
|
Zhangheng
|
78b6a4fabf
|
[UnifiedTree]: Clean up some unused dead code. (#28389)
|
2026-06-16 21:53:35 +08:00 |
|
Xinyuan Tong
|
0cb6183432
|
docs(cookbook): add GLM-5.2 deployment cookbook (#28437)
|
2026-06-16 21:49:25 +08:00 |
|