Commit Graph
14 Commits
Author SHA1 Message Date
Minglei Zhu 8a821af793 fallback to triton mm_persistent kernel when deepGemm fail (#12911) 2025-11-08 23:42:26 -08:00
Minglei Zhu c14cc47e39 [Deterministic] Optimize bmm_batch_invariant op (#12522) 2025-11-04 00:33:31 -08:00
Minglei Zhu 229256c505 [Deterministic] add deepseek v3 deterministic inference CI test (#12412) 2025-11-01 18:10:32 -07:00
Minglei Zhu e39628fd07 [2/2] Deepseek deterministic: support deepseek v3 deterministic inference on 8 x H200 (#12095) 2025-10-29 11:49:04 -07:00
Minglei Zhu f4b78d137c [1/2] deepseek deterministic: support deterministic inference for deepseek arch models on a single GPU (#12000) 2025-10-24 15:17:28 -07:00
Minglei Zhu 200a3c0bb1 [Documentation] add doc for deterministic inference (#11956) 2025-10-22 12:36:15 -05:00
Minglei Zhu f4488e9dd9 set default attention backend for deterministic inference (#11801) 2025-10-18 00:01:24 -07:00
Minglei ZhuandBaizhou Zhang 13219e1e48 completely remove mixed mode deterministic test as prefix mode could cover it (#11783)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-10-17 17:46:03 -07:00
Minglei Zhu 46ccbed2cd update GLM nightly test threshold (#10331) 2025-09-11 14:54:58 -07:00
Minglei Zhu 6ee6619b7a add zai-org/GLM-4.5-Air-FP8 model into nightly CI (#8894) 2025-08-08 01:44:19 -07:00
2ae95d17e8 Disable tp for shared experts under expert parallelism for GLM4.5 model (#8647) (#8647)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2025-08-01 12:02:35 -07:00
Minglei Zhu 25f73c6cf3 fix GLM4_MOE launch with compressed_tensor quant model (#8456) 2025-07-28 01:31:20 -07:00
Minglei Zhu 8a32355704 Feat: Support Granite 3.0 MoE in SGLang (#7959) 2025-07-17 20:56:03 -07:00
Minglei Zhu 79961afa82 optimize pad operations in fa3 to accelarate 100+us (#6077) 2025-05-07 23:40:08 -07:00