Logo
Explore Help
Register Sign In
minke.yu/sglang
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions 13 Packages Projects Releases Wiki Activity
16,048 Commits 3 Branches 0 Tags
38dc2d6cf8a69a4ff1742fe445516259ab60145f
Commit Graph
8 Commits
This Branch
This Branch
All Branches
Author SHA1 Message Date
Jackey HuaandClaude Opus 5 00a219f6c9 [Quant] Keep the flashinfer_deepgemm FP8 GEMM to 1 <= M < 32 (#32843)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 00:36:13 +08:00
Jackey HuaandClaude Opus 5 9a0bd24bed model: serve bare Qwen3Model backbone natively as an embedding model (#32457)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:49:58 +08:00
Jackey HuaandClaude Opus 4.8 ee77a7d330 [Fix] DSA: size cudagraph page_table to req_to_token width (#29379)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 12:17:34 -07:00
Jackey Hua f444b5897b [Spec][1/N] Decoupled speculative decoding: IPC protocol + cross-process request id + server flags (#27634) 2026-06-23 17:04:45 -07:00
Jackey Hua 24c5d76f74 fix: per-sequence last-token embedding in EAGLE3/MTP draft for batched multimodal spec decoding (#27846) 2026-06-11 13:33:16 -07:00
Jackey Hua 465abadd3c Add fused moe triton config for Qwen3.5-397B-A17B-FP8 (#23682) 2026-04-24 18:35:32 -07:00
jackey hua 922fbc21e2 [Perf] Tune MiniMax M2 fused moe kernel on H100 GPU (#18851) 2026-02-15 15:30:52 +08:00
jackey hua 0998de088b [Perf] Tune Llama-4-Scout-17B-16E-Instruct fused moe kernel (#17891) 2026-01-28 14:06:46 -08:00
Powered by Gitea Version: 1.27.3 Page: 343ms Template: 5ms
GitHub Default Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API