18 Commits
Author SHA1 Message Date
Yuzhen ZhouandJiajun Li 8a87079dbb Fix stale GLM MoE routing after runtime weight updates (#35883)
Co-authored-by: Jiajun Li <jiajun.li@radixark.ai>
2026-08-30 14:13:34 -07:00
Yuzhen Zhou 2d76d537e5 feat: support deterministic FA4 for GLM-4.7-Flash (#33945) 2026-08-12 16:57:34 +08:00
Yuzhen Zhou 7120f3ee13 fix: support FA4 backend for GLM4.7-flash (#33436) 2026-08-09 20:05:38 +08:00
Yuzhen Zhou b954e9cf3d [6/6][kimi-deterministic] Use deterministic seeded coins for EAGLE rejection sampling (#30822) 2026-07-24 02:11:21 -07:00
Yuzhen Zhou eb242b6c03 Seed the GDN CuteDSL correctness test inputs to fix flakiness (#32126) 2026-07-22 21:30:00 -07:00
7a973c03a0 [Bugfix] Stamp capture-time num_tokens_per_req in multi-layer EAGLE; close jit_kernel CI filter gaps (#31367)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-07-15 15:24:32 -07:00
Yuzhen Zhou 99b8f36cb1 Skip custom all-reduce v2 CUDA graph capture with torch memory saver. (#27948) 2026-06-30 14:53:29 -07:00
Yuzhen Zhou 171037c3e7 Fix Qwen3.5 deterministic batch-invariant logprobs (#27869) 2026-06-13 23:23:06 -07:00
Yuzhen ZhouandJiajun Li e03dfa8182 [3/N][Sync sglang-miles] TITO Support (#23751)
Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com>
2026-06-03 21:45:33 -04:00
Yuzhen Zhou dac78768f0 [RL][TITO] Preserve whitespace in reasoning parser outputs (#24251) 2026-05-20 19:45:09 +00:00
Yuzhen ZhouandByron Hsu 4a279d9c36 [R3] Avoid implicit CUDA sync in routed experts DP slicing (#24550)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
2026-05-06 18:37:36 -07:00
Yuzhen Zhou 6b876a7710 [ROCM][RL] Shuffle Weight In-Place to Preserve Parameter Attributes (#21825) 2026-04-02 23:43:55 -07:00
Yuzhen Zhou b719219de9 [ROCm] Use unreg path for aiter custom all-reduce during CUDA graph capture (#20155) 2026-03-09 01:09:04 -07:00
Yuzhen Zhou 63003a39cf [BUG] Support tuple hidden_states from fused MXFP4/FP8 quantization (#19643) 2026-03-02 20:39:06 -08:00
Yuzhen Zhou 4f3dc1ef5b [ROCm] Use unreg path for custom all-reduce during CUDA graph capture (#19162) 2026-02-22 23:27:31 -08:00
Yuzhen Zhou 2169025b77 turn off dit_layerwise_offload for wan on rocm (#17569) 2026-01-23 15:22:42 +08:00
4bf06635fc [diffusion] multi-platform: support diffusion on amd and fix encoder loading on MI325 (#13760)
Co-authored-by: Sabre Shao <sabre.shao@amd.com>
Co-authored-by: Yusheng (Ethan) Su <yushengsu.thu@gmail.com>
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
Co-authored-by: xsun <sunxiao04@gmail.com>
2025-12-19 15:38:46 +08:00
Yuzhen Zhou 0380ca82ef Add Batch‑Invariant RMSNorm (#12144) 2025-10-28 21:05:57 -07:00