Lucia Fang
|
99f5a6f46b
|
[flashinfer] Pass window_left at plan time for the SWA paged prefill wrapper (#31501)
|
2026-07-18 13:33:24 -07:00 |
|
Lucia Fang
|
51c5ddbe65
|
[eplb] chunk expert-weight P2P on CUDA to prevent NCCL rebalance hang (#30829)
|
2026-07-10 21:45:17 -07:00 |
|
Lucia Fang
|
1c8551169d
|
Add opt-in CUDA-graph capture-trace export (#28551)
|
2026-06-18 18:51:14 -07:00 |
|
Lucia Fang
|
05de73efd1
|
[core/model] Use explicit model arch for Llama4 attention backend auto-selection (#24232)
|
2026-05-01 15:49:30 -07:00 |
|
Lucia Fang
|
b58fa60a1f
|
[core/attention] Add SGLANG_FLASHINFER_USE_PAGED env to force paged wrapper (#24165)
|
2026-05-01 12:52:46 -07:00 |
|