8 Commits
Author SHA1 Message Date
Oguz Ulgen f15748d965 [bench] Support real-traffic replay with early-stop-aware steady-state metrics in bench_one_batch_server (#37469) 2026-09-02 17:20:25 -07:00
Oguz Ulgen d249672ad3 Scatter mm embeddings with row index_copy_ instead of masked_scatter_ to cut transient GPU memory (#37070) 2026-08-30 00:15:39 -07:00
773faf992d Reserve multimodal runtime allocations and keep padded inputs aligned (#34141)
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: wangwenchen0407 <wangwenchen@meta.com>
Co-authored-by: Hanming Lu <hanminglu@meta.com>
2026-08-12 11:04:07 -07:00
Oguz UlgenandYinghai Lu 7f6b4cb94b Add CUDA VMM multimodal feature transport (#33899)
Co-authored-by: Yinghai Lu <yinghai@meta.com>
2026-08-07 13:39:54 -07:00
Oguz UlgenandLianmin Zheng b38caebf09 [AMD] Enable gfx1250 sgl-kernel builds (#32466)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-08-06 15:47:17 -07:00
Oguz UlgenandCheng Wan c113ead98a Bump helion version to 1.4 (#32562)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2026-08-03 19:10:31 -07:00
Oguz Ulgen 3af991fb3e [AMD] Make breakable CUDA graph run on ROCm/HIP (#28173) 2026-06-19 07:16:00 -07:00
Oguz Ulgen 949326d922 Add SGLANG_ENABLE_WAR_BARRIER to force-enable the overlap scheduler WAR barrier on non-CUDA (e.g. AMD) (#27967) 2026-06-11 15:38:37 -07:00