 Артем Савкинandronnie_zheng
|
bd523dd60d
|
[NPU] [Bugfix] [Diffusion] Fixed gray images at the generation output (#23266)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-25 10:20:38 +03:00 |
|
Yujing
|
6175946db7
|
[Feature]Add MSProbe dump support in SGLang (#18349)
|
2026-04-25 10:12:50 +03:00 |
|
 Yujun Dongandhzh0425
|
21835fb0af
|
[HiCache] Prevent move_hybrid_indices from polluting radix-tree node host state (#23427)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-04-25 14:27:42 +08:00 |
|
  
|
82254bd9c5
|
[JIT Kernel] Reland JIT activation (#22094)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-24 23:00:28 -07:00 |
|
 ishandhananiandBaizhou Zhang
|
0d224e5053
|
update: b300 container for dsv4 (#23697)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-04-24 22:59:20 -07:00 |
|
YC Yen-Ching Tseng
|
adc59325bc
|
[AMD] Optimize MiniMax-M2.5 - enable fused Triton kernel for FP8 KV cache write in aiter decode path (#23620)
|
2026-04-24 22:23:49 -07:00 |
|
YC Yen-Ching Tseng
|
fb272d27db
|
[AMD] Optimize MiniMax-M2.5 - use aiter biased_grouped_topk for sigmoid scoring in MoE routing (#23611)
|
2026-04-24 22:18:08 -07:00 |
|
Shenxiu Liu
|
8471c9ebe6
|
Skip torch.cuda.empty_cache() in weight update flush path (#22998)
|
2026-04-25 12:41:39 +08:00 |
|
Baizhou Zhang
|
69485a176c
|
Small udpate gb300 recipe for deepseek v4 (#23690)
|
2026-04-24 21:35:16 -07:00 |
|
fzyzcjy
|
8a395994ed
|
docs(DeepSeek-V4): mark gb300|{small,big}|{cp,pd-disagg} verified + GB300-specific fixes (#23691)
|
2026-04-25 12:21:57 +08:00 |
|
fzyzcjy
|
d2c61acf25
|
docs(DeepSeek-V4): mark b200|small|pd-disagg + h200|small|{cp,pd-disagg} verified (#23689)
|
2026-04-25 11:57:25 +08:00 |
|
fzyzcjy
|
fd401c2fb4
|
docs(DeepSeek-V4): note SGLANG_FIX_DSV4_BASE_MODEL_LOAD for base models (#23684)
|
2026-04-25 11:27:00 +08:00 |
|
 Yuhao Yangandtrangdough
|
4a3fe2a091
|
model: support parakeet nemotron encoder (#23568)
Co-authored-by: trangdough <trangtdo22@gmail.com>
|
2026-04-25 11:00:23 +08:00 |
|
Jackey Hua
|
465abadd3c
|
Add fused moe triton config for Qwen3.5-397B-A17B-FP8 (#23682)
|
2026-04-24 18:35:32 -07:00 |
|
Lianmin Zheng
|
a4facdf3f6
|
[CI] Refactor ci_install_dependency.sh into standalone functions (#23592)
|
2026-04-24 17:39:39 -07:00 |
|
shuwenn
|
f30a6f4d7e
|
[DOC] Add DFLASH speculative decoding documentation (#23553)
|
2026-04-24 17:18:46 -07:00 |
|
Xinyi Song
|
76da28f6d6
|
[AMD][bugfix] add gate rocm >= 7.2 for bpreshuffle (#23671)
|
2026-04-24 13:26:16 -07:00 |
|
jhchouuu
|
f7e840682c
|
[AMD][MoRI] bump MoRI to v1.1.1 (#23642)
|
2026-04-24 13:12:20 -07:00 |
|
 Jia GuoandClaude Opus 4.6
|
587fd15bd2
|
perf: eliminate attention DtoD copy by passing pre-allocated output to FA (#21985)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-24 12:05:16 -07:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
Lianmin Zheng
|
6344b546c8
|
Deprecate --collect-tokens-histogram, auto-collect with --enable-metrics (#23595)
|
2026-04-24 12:00:16 -07:00 |
|
Mick
|
05696527ea
|
[diffusion] feat: support LoRA for LTX2.3 (#23649)
|
2026-04-25 01:52:41 +08:00 |
|
 
|
baa0aa670f
|
[HiCache & HybridModel] 3FS backend support DSA & mamba model (#23241)
Co-authored-by: 墨已 <kangyifei.kyf@alibaba-inc.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-04-25 00:48:01 +08:00 |
|
Kangrui Du
|
92d262f710
|
[diffusion] RL: add per-step rollout options for SDE and trajectory capture (#23151)
|
2026-04-24 23:26:16 +08:00 |
|
Siju Samuel
|
bca3dd958a
|
[Intel GPU] Enable pipeline parallelism on XPU (#23645)
|
2026-04-24 19:52:44 +08:00 |
|
 Yuwei AnandClaude Opus 4.6
|
60bbb800db
|
[Experimental] Breakable Piecewise Cuda Graph (#22218)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-24 04:33:05 -07:00 |
|
Mick
|
b3b03369a5
|
[diffusion] fix: unify LTX-2.3 HQ codepath gates for all LTX-2.3 variants (#23624)
|
2026-04-24 17:44:08 +08:00 |
|
 YC Yen-Ching Tsengandbingxche
|
b060a5ccfd
|
[AMD] Fix nightly version tag selection (#23644)
Co-authored-by: bingxche <bingxche@amd.com>
|
2026-04-24 17:39:47 +08:00 |
|
Shangming Cai
|
b8d883398d
|
Revert "[Intel GPU] Enable pipeline parallelism on XPU" (#23641)
|
2026-04-24 17:36:35 +08:00 |
|
 Ziang LiandBrayden Zhong
|
1758856762
|
[CI] Fix mxfp8 TrtllmGenMoe test (#23125)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-04-24 09:02:11 +00:00 |
|
fzyzcjy
|
92bb5c6bbe
|
Update pro fp8 checkpoint in DeepSeek V4 cookbook (#23634)
|
2026-04-24 15:58:04 +08:00 |
|
fzyzcjy
|
3a620cb761
|
Again update DeepSeek V4 cookbook (#23622)
|
2026-04-24 15:12:35 +08:00 |
|
zijiexia
|
1a37e57fb1
|
[codex] docs: note H200 DeepSeek-V4 checkpoint (#23628)
|
2026-04-24 00:06:30 -07:00 |
|
Hubert Lu
|
4cb0c4e1f3
|
[AMD] Fix memory access fault when --page-size > 1 with speculative decoding on AMD GPUs (#23596)
|
2026-04-23 23:56:36 -07:00 |
|
Mick
|
cd1fa7506a
|
[diffusion] model: support LTX2.3 high quality pipeline (#23366)
|
2026-04-24 14:18:20 +08:00 |
|
fzyzcjy
|
734e1e2965
|
Further update Deepseek V4 docs (#23617)
|
2026-04-24 13:23:50 +08:00 |
|
 
|
492883c8ca
|
Add DeepSeek V4 cookbook (#23605)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-04-23 22:10:28 -07:00 |
|
 YC Yen-Ching Tsengandbingxche
|
30909cbeeb
|
[AMD] upd local registry address (#23607)
Co-authored-by: bingxche <bingxche@amd.com>
|
2026-04-24 12:07:09 +08:00 |
|
Shaojun Zhou
|
59724e90a9
|
model: support Moss-VL (#23454)
|
2026-04-24 11:14:29 +08:00 |
|
 Siju SamuelandShangming Cai
|
bf98eb3ab7
|
[Intel GPU] Enable pipeline parallelism on XPU (#23472)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-24 10:41:51 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
b35213be11
|
[MUSA][16/N] Add MUSA backend support for layers and DeepSeek models (V2/V3/R1) (#22774)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-23 18:59:51 -07:00 |
|
Zaili Wang
|
cbc2bee547
|
[Intel CPU/XPU] SGL doc updates (#23547)
merge this one as doc change only.
|
2026-04-24 09:23:27 +08:00 |
|
Ma Mingfei
|
23e4d381f0
|
[CPU] remove RECORD_FUNCTION (#23528)
|
2026-04-24 09:18:29 +08:00 |
|
R0CKSTAR
|
87e50f20f6
|
[Apple Silicon][MLX] Cache seq_lens-derived tensors in BatchedDecodeContext (#23470)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-04-23 18:12:26 -07:00 |
|
 MARATRIXandAlex Nails
|
74c2e5bacd
|
[MUSA][8/N] Port CUDA kernels that are compatible with MUSA (#17946)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-04-23 18:04:58 -07:00 |
|
Mick
|
c0166355ae
|
[diffusion] CI: minor refactor CI (#23576)
|
2026-04-24 08:48:31 +08:00 |
|
Cheng Wan
|
d9c72bdd2b
|
Skip unselected experts in flashinfer_trtllm (#23493)
|
2026-04-23 17:30:19 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
000a2525e1
|
Move expert_mask_gpu from FusedMoE layer to StandardDispatcher (#23585)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-23 17:17:27 -07:00 |
|
Lianmin Zheng
|
95d021b523
|
Pre-set SWA cache location in CudaGraphRunner (#23552)
|
2026-04-23 16:51:29 -07:00 |
|
Lianmin Zheng
|
bb962b0046
|
Fix MoE no_combine: skip router weight in down projection (#23545)
|
2026-04-23 16:47:58 -07:00 |
|