 hjzhangandhjzhang
|
dde03d7c4a
|
[JIT] Restore the previous division behavior in per-token group quantization (#32616)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
|
2026-07-28 17:56:36 +08:00 |
|
 
|
db7e6807de
|
[BugFix] Preserve tokenizer worker fanout when skip_tokenizer_init is enabled (#30682)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-07-16 01:29:46 -07:00 |
|
 
|
c1b5c7e499
|
Fix DeepSeek V4 PP HiCache SWA allocation and layer mapping (#29106)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-06-27 22:19:14 +08:00 |
|
 
|
6779ca8d7f
|
Fix Qwen MoE precision issue with PP and all-reduce fusion (#28619)
Co-authored-by: hjzhang <zhanghjzzz@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-06-22 08:20:16 +08:00 |
|