Cheng Wan
|
582389cec5
|
[Fix] Keep diffusion encoder TP context bindings consistent (#40646)
|
2026-09-21 17:23:15 -07:00 |
|
 Tao LiandXiaoyu Zhang
|
c53cc8e1eb
|
[NPU][BugFix] Avoid M-RoPE recompilation for variable sequence lengths (#40371)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-22 08:19:25 +08:00 |
|
avalliappan-nvidia
|
61d0cf2074
|
[Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts (#32673)
|
2026-09-22 08:17:57 +08:00 |
|
 Vedant V JhaveriandCopilot
|
9fdb71732a
|
Avoid materializing GDN QKV tensors during target verification (#33778)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
2026-09-21 17:04:18 -07:00 |
|
Cheng Wan
|
506698761d
|
[unified-memory] Hierarchical cache for every unified pool shape (#37507)
|
2026-09-21 16:50:37 -07:00 |
|
Cheng Wan
|
22587fb15c
|
[Fix] Run KV canary hooks for context-parallel prefill (#40642)
|
2026-09-21 16:46:22 -07:00 |
|
 Yuxuan ZhangandXinyuan Tong
|
00986c81be
|
Support GLM-5.3-Flash hybrid attention CPU offload and PD index mapping (#40310)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-09-21 16:03:03 -07:00 |
|
YAMY
|
0229025127
|
[Spec][PP] Launch extend microbatches before the spec output exchange (#40499)
|
2026-09-21 15:47:03 -07:00 |
|
RuibinCheung
|
0c53fec476
|
[ROCm] fix: remove extra bf16 -> fp32 cast in jit grouped topk kernel path (#39775)
|
2026-09-21 15:43:21 -07:00 |
|
Kan Wu
|
d47b8c454c
|
[sgl-router] Release cancelled circuit-breaker probes (#40603)
|
2026-09-21 15:39:50 -07:00 |
|
Liangsheng Yin
|
a5c2cc517c
|
[CI] Split the CI control labels into four axes and resolve them live (#40527)
|
2026-09-21 15:37:28 -07:00 |
|
Zhang, Jiejing
|
66f19f5c46
|
[AMD] Enable HiCache for GLM-5.2 MI355X throughput recipe (#40570)
|
2026-09-21 15:28:27 -07:00 |
|
   
|
8bde82c0ad
|
[AMD] [GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP (#39339)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-21 15:24:44 -07:00 |
|
 
|
2261c2e618
|
Add MiMo-V2.6 cookbook (#40622)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
|
2026-09-21 14:58:53 -07:00 |
|
Cheng Wan
|
acac4dd9d9
|
[Refactor] Clean up parallel runtime comments (#40632)
|
2026-09-21 14:32:22 -07:00 |
|
William Hu
|
f532ad1f9a
|
Fix GLM-5.3 forget-gate shape for nvCUTEDSL verify (#40607)
|
2026-09-21 13:52:26 -07:00 |
|
jacky.cheng
|
e0c2e8dc4d
|
[AMD] Tune Qwen3.5 TP4 GDN recurrent launch on gfx950 (#39987)
|
2026-09-21 13:20:54 -07:00 |
|
Liangsheng Yin
|
1ed6822039
|
[Test] Anchor basic_perf thresholds to each metric's measured spread (#40617)
|
2026-09-21 13:02:10 -07:00 |
|
Liangsheng Yin
|
b18ca9ca44
|
[CI] Bump sgl-eval to 0.1.2 (#40620)
|
2026-09-21 13:00:54 -07:00 |
|
  
|
11e661fd45
|
[Fix] Don't free the multi-CTAs KV counter the decode graphs captured (#39175)
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-21 12:58:48 -07:00 |
|
Cheng Wan
|
44bdf225d8
|
Fix lint failure from MXFP8 reserved-slot test location (#40618)
|
2026-09-21 12:31:11 -07:00 |
|
Cheng Wan
|
bccf691b22
|
Bringing the parallel runtime up becomes a phase, not a side effect (#40345)
|
2026-09-21 12:29:50 -07:00 |
|
Cheng Wan
|
1d3243d05f
|
Take the parallel getters off the package's public surface (#40344)
|
2026-09-21 12:27:50 -07:00 |
|
Cheng Wan
|
970e946e4f
|
Retire the per-runner parallel record (#40343)
|
2026-09-21 12:26:40 -07:00 |
|
Cheng Wan
|
73f071db52
|
Deprecate the parallel getters the context answers, and ratchet them shut (#40342)
|
2026-09-21 12:25:32 -07:00 |
|
Cheng Wan
|
65be3fa71a
|
A runner and the objects it builds freeze the placement they describe (#40341)
|
2026-09-21 12:24:17 -07:00 |
|
Cheng Wan
|
2d0e94e3a3
|
Check the topology identities where the layout is written, and build at the published widths (#40340)
|
2026-09-21 12:22:59 -07:00 |
|
 ishandhananiandKangyan-Zhou
|
d5fdab7022
|
chore: add NIXL owners and CI access (#40602)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-09-21 12:21:30 -07:00 |
|
Cheng Wan
|
0db1a93adb
|
State the draft's whole topology in its scope, and read the rest from the context (#40339)
|
2026-09-21 12:19:38 -07:00 |
|
  
|
ae7a516ba7
|
feat: use XGrammar V4.1 DSML parameter constraints (#39026)
Co-authored-by: yuchuan <yuchuan.7streams@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-21 12:12:28 -07:00 |
|
Faradawn Yang
|
f0940fe3a6
|
Update DeepSeek-V4 Pro for B200 FP4 agentic PD disaggregation (#40610)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
|
2026-09-21 12:10:03 -07:00 |
|
jacky.cheng
|
90b3f8544c
|
[AMD] Use Triton softmax routing for Qwen3.5 on gfx950 (#39986)
|
2026-09-21 12:02:29 -07:00 |
|
ronnie_zheng
|
f702a0be29
|
Revert "[Diffusion] migrate the whole _register_configs from registry.py to the model own config file" (#40611)
|
2026-09-21 21:05:19 +03:00 |
|
Bingxu Chen
|
632919e498
|
[AMD] Fix DeepSeek-R1-MXFP4 accuracy with AITER FP8 (#37762)
|
2026-09-21 10:51:37 -07:00 |
|
ronnie_zheng
|
e6931ca889
|
[Diffusion] migrate the whole _register_configs from registry.py to the model own config file (#40475)
|
2026-09-21 20:45:53 +03:00 |
|
cctry
|
7a6191c4b9
|
Preallocate HiCache MHA staging before post-capture KV sizing (#40256)
|
2026-09-21 10:44:29 -07:00 |
|
cctry
|
7ad55e4386
|
[HiCache] TMA-staged host<->device KV transfer kernel (sm_90+) (#40278)
|
2026-09-21 10:38:23 -07:00 |
|
William Hu
|
0cb37c018c
|
[KDA] Enable ReplaySSM for GLM-5.3 Flash (#40517)
|
2026-09-21 10:34:08 -07:00 |
|
 Eric.Chin.AMDandThomas Wang
|
3c71bb018a
|
[AMD] Enable GLM DSA prefill top-k to the v2 kernel (#37889)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-09-21 10:31:07 -07:00 |
|
  
|
5a6a1bb883
|
[mxfp8-kv] Skip writes to the reserved CUDA-graph padding slot (#35351)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Sam Shleifer <sam@thinkingmachines.ai>
|
2026-09-22 01:21:25 +08:00 |
|
 Kan WuandShangming Cai
|
008470abd8
|
[sgl-router] Bound streaming lifetimes and release guards on idle disconnect (#40391)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-09-21 10:07:07 -07:00 |
|
Liangsheng Yin
|
800613a74b
|
[Test] Split the serving perf tests by topic into basic_perf/ and route their thresholds through a kit (#40505)
|
2026-09-21 10:05:57 -07:00 |
|
Sage
|
14e9c40a72
|
[Observability] Expose python/rust frontend identity in /server_info (#39993)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
|
2026-09-21 23:32:52 +08:00 |
|
 
|
50ec9702d0
|
[diffusion] docs: update ComfyUI sections, trimmed examples, and the RTX 5090 DiT-resident recipe (1.42x) for Qwen-Image-2.1 cookbook (#40573)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-21 21:23:11 +08:00 |
|
iridiumine
|
69d1e5cfe0
|
[Docs][NPU] Add MiMo-V2.5-Pro FP4 DFlash best practice on Ascend NPU (#40577)
|
2026-09-21 20:32:52 +08:00 |
|
amote-i
|
b410010087
|
[NPU] [DOC] Add kimi k3 cookbook for 950PR/DT Series (#40575)
|
2026-09-21 20:18:03 +08:00 |
|
Kan Wu
|
0f6761b54f
|
[sgl-router] Add SGLang-compatible DeepSeek V4.1 Flash rendering (#40532)
|
2026-09-21 18:35:09 +08:00 |
|
 
|
0abb251a20
|
[sgl-router] Match DeepSeek V4 rendering to SGLang (#40530)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-09-21 17:56:41 +08:00 |
|
 Kan WuandClaude Fable 5.1
|
2016f5e7a1
|
[sgl-router] Add Kimi-K3 rendering with SGLang parity (#40390)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
2026-09-21 17:38:49 +08:00 |
|
Zhaoyi Li
|
b86a30afba
|
[AMD][DI][CI] Add a SPUR cluster profile to AMD DI CI (#40113)
|
2026-09-21 01:56:46 -07:00 |
|