Minglei Zhu
|
200a3c0bb1
|
[Documentation] add doc for deterministic inference (#11956)
|
2025-10-22 12:36:15 -05:00 |
|
Minglei Zhu
|
f4488e9dd9
|
set default attention backend for deterministic inference (#11801)
|
2025-10-18 00:01:24 -07:00 |
|
 Minglei ZhuandBaizhou Zhang
|
13219e1e48
|
completely remove mixed mode deterministic test as prefix mode could cover it (#11783)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-10-17 17:46:03 -07:00 |
|
Minglei Zhu
|
46ccbed2cd
|
update GLM nightly test threshold (#10331)
|
2025-09-11 14:54:58 -07:00 |
|
Minglei Zhu
|
6ee6619b7a
|
add zai-org/GLM-4.5-Air-FP8 model into nightly CI (#8894)
|
2025-08-08 01:44:19 -07:00 |
|
 
|
2ae95d17e8
|
Disable tp for shared experts under expert parallelism for GLM4.5 model (#8647) (#8647)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-08-01 12:02:35 -07:00 |
|
Minglei Zhu
|
25f73c6cf3
|
fix GLM4_MOE launch with compressed_tensor quant model (#8456)
|
2025-07-28 01:31:20 -07:00 |
|
Minglei Zhu
|
8a32355704
|
Feat: Support Granite 3.0 MoE in SGLang (#7959)
|
2025-07-17 20:56:03 -07:00 |
|
Minglei Zhu
|
79961afa82
|
optimize pad operations in fa3 to accelarate 100+us (#6077)
|
2025-05-07 23:40:08 -07:00 |
|