Kurt Shuster
|
32d9fe5a32
|
[lora] Speedup triton backend sgemm calls with better grid (#22386)
|
2026-04-15 13:47:07 -07:00 |
|
 Kurt ShusterandYusheng Su
|
ff13dfee45
|
[lora][moe] Virtual experts for LoRA MoE (#22122)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
|
2026-04-13 21:19:30 +00:00 |
|
 Kurt ShusterandYusheng Su
|
f81b6df3a3
|
[lora] Fix partial MoE rank loading, VL lm_head, strict loading, deepseek on-demand (#21864)
Co-authored-by: Yusheng Su <yushengsu.thu@gmail.com>
|
2026-04-12 16:25:02 -07:00 |
|
Kurt Shuster
|
0e0091c6c8
|
[server] Add --quantization unquant to explicitly opt out of quantization (#21863)
|
2026-04-12 02:17:22 -07:00 |
|
Kurt Shuster
|
8da1cfb30d
|
[lora][moe] Decoupled LoRA MoE backend with Marlin support (#21858)
|
2026-04-11 14:59:27 -07:00 |
|
Kurt Shuster
|
db30a63a13
|
[sgl-kernel] support > 1024 experts in moe_align_block_size kernel (#21610)
|
2026-04-08 11:45:13 -07:00 |
|