 Yuzhen ZhouandJiajun Li
|
8a87079dbb
|
Fix stale GLM MoE routing after runtime weight updates (#35883)
Co-authored-by: Jiajun Li <jiajun.li@radixark.ai>
|
2026-08-30 14:13:34 -07:00 |
|
Yuzhen Zhou
|
2d76d537e5
|
feat: support deterministic FA4 for GLM-4.7-Flash (#33945)
|
2026-08-12 16:57:34 +08:00 |
|
Yuzhen Zhou
|
7120f3ee13
|
fix: support FA4 backend for GLM4.7-flash (#33436)
|
2026-08-09 20:05:38 +08:00 |
|
Yuzhen Zhou
|
b954e9cf3d
|
[6/6][kimi-deterministic] Use deterministic seeded coins for EAGLE rejection sampling (#30822)
|
2026-07-24 02:11:21 -07:00 |
|
Yuzhen Zhou
|
eb242b6c03
|
Seed the GDN CuteDSL correctness test inputs to fix flakiness (#32126)
|
2026-07-22 21:30:00 -07:00 |
|
 
|
7a973c03a0
|
[Bugfix] Stamp capture-time num_tokens_per_req in multi-layer EAGLE; close jit_kernel CI filter gaps (#31367)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-07-15 15:24:32 -07:00 |
|
Yuzhen Zhou
|
99b8f36cb1
|
Skip custom all-reduce v2 CUDA graph capture with torch memory saver. (#27948)
|
2026-06-30 14:53:29 -07:00 |
|
Yuzhen Zhou
|
171037c3e7
|
Fix Qwen3.5 deterministic batch-invariant logprobs (#27869)
|
2026-06-13 23:23:06 -07:00 |
|
 Yuzhen ZhouandJiajun Li
|
e03dfa8182
|
[3/N][Sync sglang-miles] TITO Support (#23751)
Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com>
|
2026-06-03 21:45:33 -04:00 |
|
Yuzhen Zhou
|
dac78768f0
|
[RL][TITO] Preserve whitespace in reasoning parser outputs (#24251)
|
2026-05-20 19:45:09 +00:00 |
|
 Yuzhen ZhouandByron Hsu
|
4a279d9c36
|
[R3] Avoid implicit CUDA sync in routed experts DP slicing (#24550)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2026-05-06 18:37:36 -07:00 |
|
Yuzhen Zhou
|
6b876a7710
|
[ROCM][RL] Shuffle Weight In-Place to Preserve Parameter Attributes (#21825)
|
2026-04-02 23:43:55 -07:00 |
|
Yuzhen Zhou
|
b719219de9
|
[ROCm] Use unreg path for aiter custom all-reduce during CUDA graph capture (#20155)
|
2026-03-09 01:09:04 -07:00 |
|
Yuzhen Zhou
|
63003a39cf
|
[BUG] Support tuple hidden_states from fused MXFP4/FP8 quantization (#19643)
|
2026-03-02 20:39:06 -08:00 |
|
Yuzhen Zhou
|
4f3dc1ef5b
|
[ROCm] Use unreg path for custom all-reduce during CUDA graph capture (#19162)
|
2026-02-22 23:27:31 -08:00 |
|
Yuzhen Zhou
|
2169025b77
|
turn off dit_layerwise_offload for wan on rocm (#17569)
|
2026-01-23 15:22:42 +08:00 |
|
   
|
4bf06635fc
|
[diffusion] multi-platform: support diffusion on amd and fix encoder loading on MI325 (#13760)
Co-authored-by: Sabre Shao <sabre.shao@amd.com>
Co-authored-by: Yusheng (Ethan) Su <yushengsu.thu@gmail.com>
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
Co-authored-by: xsun <sunxiao04@gmail.com>
|
2025-12-19 15:38:46 +08:00 |
|
Yuzhen Zhou
|
0380ca82ef
|
Add Batch‑Invariant RMSNorm (#12144)
|
2025-10-28 21:05:57 -07:00 |
|