Ethan (Yusheng) Su
|
3c5f115741
|
Split #32584 into 2/2: [LoRA] Shard attention LoRA by attn-TP and allow dynamic LoRA with dp attention (#32708)
|
2026-07-31 15:37:13 -07:00 |
|
Ethan (Yusheng) Su
|
7e996a5d0d
|
Split #32584 into 1/2: [LoRA] Guard DP-attention idle forwards against stale LoRA batch state (#32707)
|
2026-07-31 14:23:01 -07:00 |
|
Ethan (Yusheng) Su
|
0a49226d19
|
[LoRA] 1/n Per-rank tensor serialization for load_lora_adapter_from_tensors under dp_size > 1 (#32580)
|
2026-07-28 14:28:56 -07:00 |
|
Ethan (Yusheng) Su
|
ee1736f39a
|
[LoRA] Support LoRA under the breakable/full prefill CUDA graph (#30988)
|
2026-07-26 22:10:03 -07:00 |
|
Ethan (Yusheng) Su
|
3849beb7e3
|
[CI] Fix XPU platform test on machines without the XPU sgl-kernel op (#32298)
|
2026-07-24 16:42:36 +08:00 |
|
 Ethan (Yusheng) SuandCursor
|
a24c374f84
|
[lora] Remove synchronous .any().item() guard in LoRA MoE prefill path (#25531)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-21 23:58:57 +08:00 |
|
 
|
e9dea79755
|
(3/n - prefill optimize)[LoRA][MoE] Optimize virtual experts: remove CPU-GPU sync & multi-block CUDA JIT histogram (#24262)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-11 16:36:57 -07:00 |
|