Commit Graph
593 Commits
Author SHA1 Message Date
Xiaoyu ZhangandClaude Opus 4.8 99f636a86f [Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 08:35:09 +08:00
Mohammad Miadh AngkadandBrayden Zhong 0c29c8fece Bump FlashInfer to 0.6.15.post1 (#31927)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-07-22 14:21:59 -07:00
Mohammad Miadh Angkad 3d82dacd58 Bump CuTe DSL to 4.6.0 (#31714) 2026-07-20 02:11:59 -07:00
02236fa38c Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
2026-07-19 22:57:37 -07:00
Baizhou Zhang 304a529558 Revert "Bump FlashInfer to 0.6.15 and revert regressions" (#31625) 2026-07-17 16:46:33 -07:00
Lianmin Zheng c95026aed3 Upgrade llguidance to 1.7.6 (#31484) 2026-07-17 16:31:44 -07:00
sglang-botandsglang-bot 0ad0ff2e9e chore: bump sglang-kernel version to 0.4.5 (#31618)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-17 15:58:36 -07:00
Mohammad Miadh AngkadandBrayden Zhong d67aa05697 Bump FlashInfer to 0.6.15 and revert regressions (#31502)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-07-17 00:50:12 -07:00
Khoa PhamandClaude Fable 5 dc60f65661 chore: bump tokenspeed_mla to 0.1.8 (#31385)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:37:54 -07:00
423b8485fb [Quantization] add humming quantization kernel (#23754)
Co-authored-by: guzekai01 <zekai01@antgroup.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@gmail.com>
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-07-14 08:42:56 +08:00
Mick 649ce5dd3d model: support Pi0.5 (#30633) 2026-07-11 07:50:58 +08:00
Baizhou Zhang 3f1694f5e0 Update sgl-deep-gemm to 0.1.4.post1 (#30697) 2026-07-10 13:30:43 -07:00
Xinyuan Tong b76dd0be69 Fix Mistral GSM8K chat eval (#27757) 2026-07-09 21:08:48 -07:00
2c6cd1ef41 [Dep] Upgrade flashinfer to 0.6.14 (#29910)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-07-09 17:52:33 -07:00
Chetan Kumar VermaandMa Mingfei b3ab56545b Add Accuracy Benchmark for OCR models (#25364)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-06 16:18:43 +08:00
Baizhou Zhang c312cdd3a7 Upgrading tvm-ffi/sgl-deep-gemm/tilelang (#29554) 2026-07-01 12:32:18 -07:00
Jialin Ouyang 40594bd381 [passthrough] engine: zstd request-body decompression + header overrides (#29684) 2026-06-30 23:48:56 -07:00
Mohammad Miadh Angkad bae78a44da [Deps] Bump transformers to 5.12.1 (#29393) 2026-06-30 00:54:01 -07:00
cfc0a0e0e0 Add Intel Quantization Support in SGLang (#18139)
Signed-off-by: Mengni Wang <mengni.wang@intel.com>
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Weiwei <weiwei1.zhang@intel.com>
2026-06-26 09:54:35 +08:00
Khoa PhamandCursor 0642cd5020 (chore): bump tokenspeed_mla to 0.1.7 (#28759)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-21 21:25:37 -07:00
Mick a51d56d948 CI: Pin flash-attn-4 for diffusion CI consistency (#28838) 2026-06-21 21:13:41 +08:00
Lianmin Zheng 8a3d6c3403 Sort pyproject dependency lists (#28811) 2026-06-20 17:24:57 -07:00
sglang-botandsglang-bot b1d18d562b chore: bump sglang-kernel version to 0.4.4 (#28572)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-17 16:13:39 -07:00
Baizhou Zhang 27291118b9 Upgrade sgl-deep-gemm to 0.1.3 (#28402) 2026-06-16 23:06:36 -07:00
feliang-gitandxutizhou 92b42c8d8a LPLB: linear-programming load balancer for MoE expert parallelism (#24515)
Co-authored-by: xutizhou <xutingz@nvidia.com>
2026-06-16 10:19:42 -07:00
Xinyu Zhang 1800d7caa6 Bump ray minimum version to 2.55.1 (#27724) 2026-06-12 20:49:11 -07:00
Khoa Pham a0c6e0b3a4 chore: bump tokenspeed_mla 0.1.1 -> 0.1.6 (#28116) 2026-06-12 19:57:36 -07:00
Mohammad Miadh Angkad 7f706f4cfb [Deps] Bump FI to 0.6.12 and cutedsl to 4.5.2 (#26854) 2026-06-03 12:09:18 -07:00
Baizhou Zhang 68caf49154 Update sgl-deep-gemm to 0.1.2 (#26993) 2026-06-02 13:51:54 -07:00
Mick 64a1dec8b6 [diffusion] feat: add realtime webui super resolution controls (#27026) 2026-06-02 20:29:21 +08:00
Mick 3b26644bc4 [diffusion] misc: add realtime-webui (#26959) 2026-06-02 14:13:02 +08:00
Mick 2fc548f250 [diffusion] model: support lingot-world (#26954) 2026-06-02 13:52:49 +08:00
popsiclexuandpopsiclexu 951fa05a09 [MoE] Support BF16 standard A2A with DeepGEMM runner (#26473)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
2026-06-01 20:40:38 -07:00
Liangsheng Yin ed85bcf8c3 pin kernels<0.15 (#26704) 2026-05-29 01:46:57 -07:00
Xinyuan Tong 79c844527c Upgrade xgrammar to 0.2.1 (#25676) 2026-05-29 11:40:07 +08:00
sglang-botandsglang-bot 14f81a67d9 chore: bump sglang-kernel version to 0.4.3 (#26421)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-05-27 11:31:11 -07:00
Kangyan-ZhouandLiangsheng Yin caa9f08294 [CI] Force-reinstall nvidia-cutlass-dsl-libs-cu13 last to avoid wheel-mix TypeError (#25958)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-05-21 22:01:42 +08:00
Kangyan-Zhou 4ea8282cb7 [Revert] nvidia-cutlass-dsl[cu13] 4.5.1 -> 4.5.0 (#25938) 2026-05-21 14:36:56 +08:00
Mohammad Miadh Angkad a449ee4822 [Deps] Use cu13 extra for nvidia cutlass dsl (#25576) 2026-05-21 10:31:27 +08:00
Matt Van HornandMatt Van Horn e99f87c974 fix: add missing distro dependency to runtime docker image (#25817)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
2026-05-20 00:46:32 -07:00
Xinyuan Tong aad00b0ed8 Upgrade transformers to 5.8.1 (#25451) 2026-05-19 22:20:30 +08:00
Baizhou Zhangandhnyls2002 b79e4b1e68 [Fix] Try to fix error caused by latest cutedsl packages (#25690)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-05-18 16:51:32 -07:00
0c19540550 [Fix] Fix gpt oss triton kernels and upgrade flashinfer back to 0.6.11.post1 (#25335)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: mmangkad <mmangkad@users.noreply.github.com>
2026-05-15 01:04:56 -07:00
Liangsheng Yin 22dfcdaa04 revert flashinfer 0.6.11 bumps (#25310) 2026-05-14 15:28:58 -07:00
Baizhou Zhangandpranjalssh b7f856df70 DeepSeek V4 w4a4 MegaMoE (#25052)
Co-authored-by: pranjalssh <adkz.photos@gmail.com>
2026-05-13 18:35:32 -07:00
Qiaolin Yu 7618ad7075 [attn backend] Integrate tokenspeed_mla prefill/decode kernels (fp8 kv cache, blackwell) (#24925) 2026-05-13 17:36:17 -07:00
Baizhou Zhang 51a9403104 Update flashinfer to 0.6.11.post1 (#25129) 2026-05-13 00:12:19 -07:00
Brayden Zhongandb8zhong d5f3254ed1 [Dependency] Flashinfer 0.6.8post1 -> 0.6.11 (#24452)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-12 14:38:32 -07:00
+6 35870d55ac Deepseek V4 (#23882)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Qiaolin Yu <90088090+qiaolin-yu@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <11704492+yushengsu-thu@users.noreply.github.com>
Co-authored-by: Mingyi <27337995+wisclmy0611@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+againstentropy@users.noreply.github.com>
2026-05-07 18:32:21 -07:00
Baizhou Zhang ecb786c8d7 [Kernel] Deprecate DeepGemm in sgl kernel and apply custom wheel sgl-deep-gemm (#24268) 2026-05-06 18:59:01 -07:00