 sglang-botandsglang-bot
|
6252993afe
|
chore: update CI test est_time values (#38238)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-09-06 17:49:41 -07:00 |
|
   
|
15aa2fb843
|
[ROCm] Take the fused DSA metadata kernels and drop redundant work from the absorb path (#37124)
Co-authored-by: yanyuan.qin <yanyuan.qin@amd.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-09-06 17:39:37 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
30705c004c
|
[Deepseek V4] Keep fp32 routing weights in the mxfp4 trtllm MoE (#33608)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-09-07 00:12:55 +00:00 |
|
AMD-yanfeiwang
|
2e8c03e2c7
|
Fix inflated row pitch when a CP round-robin shard has a single row (#34142)
|
2026-09-06 15:40:14 -07:00 |
|
JohnQinAMD
|
2c05ed4e77
|
[ROCm] Stage large pageable H2D copies instead of pinning them in place (#37720)
|
2026-09-06 12:28:58 -07:00 |
|
 
|
31d28a2961
|
[NPU] Fix failed test cases in pr‑test‑npu and improve execution efficiency (#38112)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-09-07 01:48:27 +08:00 |
|
s
|
28457f0dca
|
fix(gpt-oss): avoid duplicate MoE reduction with DP attention (#37199)
|
2026-09-07 00:33:04 +08:00 |
|
Mick
|
f3d05644db
|
[diffusion] docs+skill: document which components to stream under layerwise offload (#35674)
|
2026-09-06 23:12:36 +08:00 |
|
 MickandClaude Fable 5
|
ade1da017f
|
[diffusion] docs: verify the DGX Spark H3 recipe (#37456)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-09-06 23:10:27 +08:00 |
|
Xiaoyu Zhang
|
be00a543a7
|
perf: use Gumbel-max trick in the main sampler to cut decode CPU dispatch (#38117)
|
2026-09-06 22:40:27 +08:00 |
|
 
|
5bebe7a033
|
[Router] Add bucket-aware policy domains and native cache indexing (#38108)
Signed-off-by: Vincent Gao <vincentbo@linux.alibaba.com>
Co-authored-by: inkcherry <mingzhi.liu@amd.com>
Co-authored-by: yangbodong22011 <13137470+yangbodong22011@users.noreply.github.com>
|
2026-09-06 19:47:51 +08:00 |
|
 MickandClaude Fable 5.1
|
a176ba2f7b
|
[diffusion] feat: measure warmup memory and layer usage per phase for residency calibration (1/4) (#37916)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-06 19:42:42 +08:00 |
|
Mick
|
938dc5621d
|
[diffusion] refactor: reuse plain state-dict loading without per-model classes (#38127)
|
2026-09-06 18:39:21 +08:00 |
|
+2        
|
97c6978369
|
GLM-5.3-Flash support (#36507)
Co-authored-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk>
Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: zanes-ops <zanes@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: Ehsan Akhgari <ehsan.akhgari@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
|
2026-09-06 02:27:59 -07:00 |
|
Mick
|
a9944aec01
|
[diffusion] CI: guard the allocated vram peak with reporting the reserved one (#38172)
|
2026-09-06 17:22:16 +08:00 |
|
 Mickandmickqian
|
8ef646a5c6
|
fix(vlm): contain EPD request lifecycle failures (#36944)
Co-authored-by: mickqian <mickqian@users.noreply.github.com>
|
2026-09-06 16:05:10 +08:00 |
|
Xiaoyu Zhang
|
d61378af77
|
docs(diffusion): add per-model tuning decision table to performance guide (#38148)
|
2026-09-06 15:26:40 +08:00 |
|
Vincent Gao
|
ae54ccb25d
|
[Router] Publish cache-aware load state (#38139)
|
2026-09-06 15:22:25 +08:00 |
|
Zhang, Jiejing
|
6cee9285a3
|
[ROCm] Make DSA indexer top-k exact with cooperative selection (#37591)
|
2026-09-05 23:55:56 -07:00 |
|
Mick
|
e3f7097591
|
[diffusion] refactor: consolidate plain state-dict component loaders (#38128)
|
2026-09-06 13:52:54 +08:00 |
|
Depend
|
67e3ccda97
|
[diffusion] fix: restore non-layer placeholders before releasing host copies (#38171)
|
2026-09-06 13:41:30 +08:00 |
|
Mick
|
febb360519
|
[VLM] retire aborted disaggregated prefill results (#36988)
|
2026-09-06 10:15:33 +08:00 |
|
Liangsheng Yin
|
f5819b09bf
|
Revert "[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting" (#38163)
|
2026-09-05 17:28:46 -07:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
09daea94ac
|
Support NoPE layers in the tokenspeed_mla FP8 prefill hook (#38152)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-05 16:51:04 -07:00 |
|
yuttian1
|
514b45fd34
|
[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting (#30315)
|
2026-09-05 16:39:40 -07:00 |
|
 Alex NailsandClaude Opus 5
|
6a0c55fd6c
|
[CI] Pin the Rust TreeCore build to the resolved libtorch instead of interpreter discovery (#37696)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-05 15:30:32 -07:00 |
|
 
|
77aee20259
|
[Model] Add support for Nanbeige4.2 (#32151)
Co-authored-by: root <lizongqiang@kanzhun.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-06 03:57:27 +08:00 |
|
 Xiaoyu ZhangandWaterpine
|
ccf9fe6590
|
[Kernel] Add KDA FP8 skinny GEMM for SM120 (#38082)
Co-authored-by: Waterpine <biansonghz@gmail.com>
|
2026-09-05 22:27:06 +08:00 |
|
Cherry_ming
|
9acbf75159
|
[CI][NPU] Fix pr-test-npu failing at Install dependencies with set: Illegal option -o pipefail (#38132)
|
2026-09-05 22:26:47 +08:00 |
|
Xiaoyu Zhang
|
dc2843801d
|
perf(lfm2): fuse gating and short convolution on SM90 (#37622)
|
2026-09-05 21:52:16 +08:00 |
|
Beihao Zhou
|
1e6f18bfeb
|
[MoE Refactor] Migrate SM100 trtllm-gen mxfp4 MoE onto MoeRunner (#32405)
|
2026-09-05 13:48:10 +00:00 |
|
Xiaoyu Zhang
|
eda10c3678
|
[Diffusion] Enable breakable CUDA graph for JoyEcho (#38110)
|
2026-09-05 21:45:53 +08:00 |
|
Mick
|
5df60a21cd
|
fix(vlm): harden EPD receiver validation and liveness (#36945)
|
2026-09-05 21:22:37 +08:00 |
|
 
|
4b802c052b
|
test(npu): remove obsolete npu pr nightly cases, move accuracy cases to full (#37990)
Co-authored-by: Sugar920 <Sugar920@users.noreply.github.com>
Co-authored-by: Claude Code <noreply@anthropic.com>
|
2026-09-05 20:58:52 +08:00 |
|
Mick
|
a18106bbc3
|
fix(vlm): make EPD cache publication transactional (#36949)
|
2026-09-05 20:33:23 +08:00 |
|
 
|
bd16c22a04
|
[diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-05 18:12:30 +08:00 |
|
 DayuxiaoshuiandXiaoyu Zhang
|
50c1bf0db0
|
[Diffusion] Port the Wan VAE decoder fast paths to the Qwen-Image VAE (#38020)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-05 18:02:33 +08:00 |
|
hhhh1252023
|
0948e6ebed
|
[CI] Remove metrics artifact mechanism from nightly NPU workflows (#35489)
|
2026-09-05 17:16:52 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
d49180019b
|
fix(moe): cast filtered-activation expert_ids to int32 for torch.compile (#38085)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
|
2026-09-05 17:13:07 +08:00 |
|
Xiaoyu Zhang
|
a74470e904
|
fix(mamba): unify causal_conv1d col* dtype to x (MiniCPM-V-4.6 GDN prefill bf16/fp16 mismatch) (#38039)
|
2026-09-05 17:08:20 +08:00 |
|
jacky.cheng
|
0bdc15d20f
|
[AMD] Fix ROCm VAE Conv2D fast path breaking spatial-parallel decode (#34424)
|
2026-09-05 01:00:25 -07:00 |
|
Xiaoyu Zhang
|
da76fa073f
|
[diffusion] fix: fix host-resident vocab tables loaded on GPU (#38012)
|
2026-09-05 14:03:21 +08:00 |
|
Mick
|
0ea8378085
|
[diffusion] feat: support request-scoped skip-softmax attention (#37959)
|
2026-09-05 13:50:17 +08:00 |
|
 Shuwen WangandClaude Opus 5
|
4b44a1cde2
|
[Refactor] Let eviction policies take construction parameters (#37795)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-09-05 13:48:30 +08:00 |
|
Xinyuan Tong
|
32a1d55431
|
fix(modelopt_fp4): skip NVFP4 swiglu-fusion interleave for shared experts with swiglu_limit (#37378)
|
2026-09-04 22:42:29 -07:00 |
|
Ke Bao
|
ae3205ba28
|
Fail fast on undersized swa pool (#37610)
|
2026-09-05 13:40:37 +08:00 |
|
 
|
09f542b23a
|
[CI] Add /rerun-test --changed to rerun every test file a PR modifies (#37618)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-04 22:38:11 -07:00 |
|
 Yash AkhauriandXiaoyu Zhang
|
756d0e0a85
|
[Model] Add K2 Horizon FP8 checkpoint support (#38033)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-05 13:18:26 +08:00 |
|
 
|
e980c1a2f1
|
fix(glm4v): disambiguate mixed image video offsets (#37971)
Co-authored-by: duxin <xinheng.dx@alibaba-inc.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-05 13:16:16 +08:00 |
|
Liangsheng Yin
|
0454c074b4
|
[mem_cache] Clean up unified allocator leftovers (#38103)
|
2026-09-04 22:02:50 -07:00 |
|