Logo
Explore Help
Register Sign In
minke.yu/sglang
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions 10 Packages Projects Releases Wiki Activity
18,691 Commits 2 Branches 0 Tags
dsv41-pd
Commit Graph
6 Commits
This Branch
This Branch
All Branches
Author SHA1 Message Date
Rahul Vijayaraghavan d4ad368ed9 [Intel XPU] Enable fused_moe_triton tuning on XPU and add tuned DeepSeek-OCR-2 configs (#28723) 2026-09-16 10:20:22 +08:00
Rahul Vijayaraghavan fa0ced195e [XPU] Enable breakable prefill CUDA graph on XPU (#30273) 2026-07-21 09:09:40 +08:00
Rahul VijayaraghavanandMa Mingfei d4963f5c55 Fix prefill CUDA graph disabled for deeply-nested multimodal models (#30006)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-08 12:11:58 +08:00
Rahul VijayaraghavanandMa Mingfei 4b5c612257 Skip redundant moe_sum_reduce for single-expert routing on XPU (#22660)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-07 09:03:16 +08:00
Rahul Vijayaraghavan 1aee04e7df [XPU] Support apply_router_weight_on_input for Llama4 for fused_experts (#22654)
merge this one as it is xpu only change.
2026-04-29 10:44:48 +08:00
Rahul Vijayaraghavan ac2819c81f Fix assertion tolerance for bf16 precision in triton attention UT (#17461)
Signed-off-by: Rahul Vijayaraghavan <rahul.vijayaraghavan@intel.com>
2026-03-03 13:43:58 -08:00
Powered by Gitea Version: 1.27.3 Page: 444ms Template: 3ms
GitHub Default Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API