Mohammad Miadh Angkad
|
3d82dacd58
|
Bump CuTe DSL to 4.6.0 (#31714)
|
2026-07-20 02:11:59 -07:00 |
|
       
|
02236fa38c
|
Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
|
2026-07-19 22:57:37 -07:00 |
|
Baizhou Zhang
|
304a529558
|
Revert "Bump FlashInfer to 0.6.15 and revert regressions" (#31625)
|
2026-07-17 16:46:33 -07:00 |
|
Lianmin Zheng
|
c95026aed3
|
Upgrade llguidance to 1.7.6 (#31484)
|
2026-07-17 16:31:44 -07:00 |
|
 sglang-botandsglang-bot
|
0ad0ff2e9e
|
chore: bump sglang-kernel version to 0.4.5 (#31618)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-17 15:58:36 -07:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
d67aa05697
|
Bump FlashInfer to 0.6.15 and revert regressions (#31502)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-07-17 00:50:12 -07:00 |
|
 Khoa PhamandClaude Fable 5
|
dc60f65661
|
chore: bump tokenspeed_mla to 0.1.8 (#31385)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 22:37:54 -07:00 |
|
     
|
423b8485fb
|
[Quantization] add humming quantization kernel (#23754)
Co-authored-by: guzekai01 <zekai01@antgroup.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@gmail.com>
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-14 08:42:56 +08:00 |
|
Mick
|
649ce5dd3d
|
model: support Pi0.5 (#30633)
|
2026-07-11 07:50:58 +08:00 |
|
Baizhou Zhang
|
3f1694f5e0
|
Update sgl-deep-gemm to 0.1.4.post1 (#30697)
|
2026-07-10 13:30:43 -07:00 |
|
Xinyuan Tong
|
b76dd0be69
|
Fix Mistral GSM8K chat eval (#27757)
|
2026-07-09 21:08:48 -07:00 |
|
 
|
2c6cd1ef41
|
[Dep] Upgrade flashinfer to 0.6.14 (#29910)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-07-09 17:52:33 -07:00 |
|
 Chetan Kumar VermaandMa Mingfei
|
b3ab56545b
|
Add Accuracy Benchmark for OCR models (#25364)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-06 16:18:43 +08:00 |
|
Baizhou Zhang
|
c312cdd3a7
|
Upgrading tvm-ffi/sgl-deep-gemm/tilelang (#29554)
|
2026-07-01 12:32:18 -07:00 |
|
Jialin Ouyang
|
40594bd381
|
[passthrough] engine: zstd request-body decompression + header overrides (#29684)
|
2026-06-30 23:48:56 -07:00 |
|
Mohammad Miadh Angkad
|
bae78a44da
|
[Deps] Bump transformers to 5.12.1 (#29393)
|
2026-06-30 00:54:01 -07:00 |
|
  
|
cfc0a0e0e0
|
Add Intel Quantization Support in SGLang (#18139)
Signed-off-by: Mengni Wang <mengni.wang@intel.com>
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Weiwei <weiwei1.zhang@intel.com>
|
2026-06-26 09:54:35 +08:00 |
|
 Khoa PhamandCursor
|
0642cd5020
|
(chore): bump tokenspeed_mla to 0.1.7 (#28759)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-21 21:25:37 -07:00 |
|
Mick
|
a51d56d948
|
CI: Pin flash-attn-4 for diffusion CI consistency (#28838)
|
2026-06-21 21:13:41 +08:00 |
|
Lianmin Zheng
|
8a3d6c3403
|
Sort pyproject dependency lists (#28811)
|
2026-06-20 17:24:57 -07:00 |
|
 sglang-botandsglang-bot
|
b1d18d562b
|
chore: bump sglang-kernel version to 0.4.4 (#28572)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-06-17 16:13:39 -07:00 |
|
Baizhou Zhang
|
27291118b9
|
Upgrade sgl-deep-gemm to 0.1.3 (#28402)
|
2026-06-16 23:06:36 -07:00 |
|
 feliang-gitandxutizhou
|
92b42c8d8a
|
LPLB: linear-programming load balancer for MoE expert parallelism (#24515)
Co-authored-by: xutizhou <xutingz@nvidia.com>
|
2026-06-16 10:19:42 -07:00 |
|
Xinyu Zhang
|
1800d7caa6
|
Bump ray minimum version to 2.55.1 (#27724)
|
2026-06-12 20:49:11 -07:00 |
|
Khoa Pham
|
a0c6e0b3a4
|
chore: bump tokenspeed_mla 0.1.1 -> 0.1.6 (#28116)
|
2026-06-12 19:57:36 -07:00 |
|
Mohammad Miadh Angkad
|
7f706f4cfb
|
[Deps] Bump FI to 0.6.12 and cutedsl to 4.5.2 (#26854)
|
2026-06-03 12:09:18 -07:00 |
|
Baizhou Zhang
|
68caf49154
|
Update sgl-deep-gemm to 0.1.2 (#26993)
|
2026-06-02 13:51:54 -07:00 |
|
Mick
|
64a1dec8b6
|
[diffusion] feat: add realtime webui super resolution controls (#27026)
|
2026-06-02 20:29:21 +08:00 |
|
Mick
|
3b26644bc4
|
[diffusion] misc: add realtime-webui (#26959)
|
2026-06-02 14:13:02 +08:00 |
|
Mick
|
2fc548f250
|
[diffusion] model: support lingot-world (#26954)
|
2026-06-02 13:52:49 +08:00 |
|
 popsiclexuandpopsiclexu
|
951fa05a09
|
[MoE] Support BF16 standard A2A with DeepGEMM runner (#26473)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
|
2026-06-01 20:40:38 -07:00 |
|
Liangsheng Yin
|
ed85bcf8c3
|
pin kernels<0.15 (#26704)
|
2026-05-29 01:46:57 -07:00 |
|
Xinyuan Tong
|
79c844527c
|
Upgrade xgrammar to 0.2.1 (#25676)
|
2026-05-29 11:40:07 +08:00 |
|
 sglang-botandsglang-bot
|
14f81a67d9
|
chore: bump sglang-kernel version to 0.4.3 (#26421)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-05-27 11:31:11 -07:00 |
|
 Kangyan-ZhouandLiangsheng Yin
|
caa9f08294
|
[CI] Force-reinstall nvidia-cutlass-dsl-libs-cu13 last to avoid wheel-mix TypeError (#25958)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-05-21 22:01:42 +08:00 |
|
Kangyan-Zhou
|
4ea8282cb7
|
[Revert] nvidia-cutlass-dsl[cu13] 4.5.1 -> 4.5.0 (#25938)
|
2026-05-21 14:36:56 +08:00 |
|
Mohammad Miadh Angkad
|
a449ee4822
|
[Deps] Use cu13 extra for nvidia cutlass dsl (#25576)
|
2026-05-21 10:31:27 +08:00 |
|
 Matt Van HornandMatt Van Horn
|
e99f87c974
|
fix: add missing distro dependency to runtime docker image (#25817)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
|
2026-05-20 00:46:32 -07:00 |
|
Xinyuan Tong
|
aad00b0ed8
|
Upgrade transformers to 5.8.1 (#25451)
|
2026-05-19 22:20:30 +08:00 |
|
 Baizhou Zhangandhnyls2002
|
b79e4b1e68
|
[Fix] Try to fix error caused by latest cutedsl packages (#25690)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-05-18 16:51:32 -07:00 |
|
  
|
0c19540550
|
[Fix] Fix gpt oss triton kernels and upgrade flashinfer back to 0.6.11.post1 (#25335)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: mmangkad <mmangkad@users.noreply.github.com>
|
2026-05-15 01:04:56 -07:00 |
|
Liangsheng Yin
|
22dfcdaa04
|
revert flashinfer 0.6.11 bumps (#25310)
|
2026-05-14 15:28:58 -07:00 |
|
 Baizhou Zhangandpranjalssh
|
b7f856df70
|
DeepSeek V4 w4a4 MegaMoE (#25052)
Co-authored-by: pranjalssh <adkz.photos@gmail.com>
|
2026-05-13 18:35:32 -07:00 |
|
Qiaolin Yu
|
7618ad7075
|
[attn backend] Integrate tokenspeed_mla prefill/decode kernels (fp8 kv cache, blackwell) (#24925)
|
2026-05-13 17:36:17 -07:00 |
|
Baizhou Zhang
|
51a9403104
|
Update flashinfer to 0.6.11.post1 (#25129)
|
2026-05-13 00:12:19 -07:00 |
|
 Brayden Zhongandb8zhong
|
d5f3254ed1
|
[Dependency] Flashinfer 0.6.8post1 -> 0.6.11 (#24452)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-12 14:38:32 -07:00 |
|
+6        
|
35870d55ac
|
Deepseek V4 (#23882)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Qiaolin Yu <90088090+qiaolin-yu@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <11704492+yushengsu-thu@users.noreply.github.com>
Co-authored-by: Mingyi <27337995+wisclmy0611@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+againstentropy@users.noreply.github.com>
|
2026-05-07 18:32:21 -07:00 |
|
Baizhou Zhang
|
ecb786c8d7
|
[Kernel] Deprecate DeepGemm in sgl kernel and apply custom wheel sgl-deep-gemm (#24268)
|
2026-05-06 18:59:01 -07:00 |
|
  
|
952b3caf18
|
feat: use structural tags to enable strict tool calling and reasoning for more models (#21722)
Signed-off-by: Yuchuan <yuchuan.7streams@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Ubospica <ubospica@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-05-04 02:30:28 -07:00 |
|
    
|
88bb5dffe4
|
[Dependency] Upgrade to Torch 2.11.0 (#21247)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-05-02 12:25:36 -07:00 |
|