Lianmin Zheng
|
67e12131df
|
Build Rust extensions on demand in source checkouts (#34994)
|
2026-08-16 14:58:06 -07:00 |
|
Baizhou Zhang
|
7769f54feb
|
[Kimi-K3] Use explicit SiTU activation for MegaMoE (#34883)
|
2026-08-15 01:41:59 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
5e65dd01a7
|
Remove the torchao integration (--torchao-config) (#34304)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-08-14 21:49:11 +08:00 |
|
Mohammad Miadh Angkad
|
00e57d74f0
|
Bump FlashInfer to 0.6.17 and remove Kimi K3 workarounds (#33997)
|
2026-08-12 02:17:26 -07:00 |
|
Mohammad Miadh Angkad
|
a3bd7d9401
|
Bump CuTeDSL to 4.6.2 (#34372)
|
2026-08-12 00:35:35 +08:00 |
|
Liangsheng Yin
|
c80a38edcd
|
[Fix] Pin cuda-tile to 1.6.0rc5 to unblock Python 3.10 x86_64 installs (#34321)
|
2026-08-10 15:09:47 -07:00 |
|
 Sam ShleiferandClaude Fable 5
|
afb4f37ca5
|
[Inkling] silu_and_mul: replace helion kernels with plain Triton (#33903)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 15:00:23 +08:00 |
|
 
|
f64328c7f6
|
[diffusion] feat: support quant-videogen prq kv-cache quantization (memory-saving) for causal-dit (#32581)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-08-08 12:56:54 +08:00 |
|
Baizhou Zhang
|
eb3cc879e0
|
Install DeepEP from release wheels (#33932)
|
2026-08-07 15:38:44 -07:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
434e646282
|
[Deps] Upgrade CUDA PyTorch stack to 2.13 (#28836)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-08-06 12:08:44 -07:00 |
|
+26        
|
abddb1c7e9
|
[Kimi] Support kimi-k3 (#32541)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai>
Co-authored-by: Zijie Xia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: zhangxiaohao <1024393531@qq.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: RolaoDenthu <xinyisong0111@gmail.com>
Co-authored-by: pigeonsoup <32922982+pigeonsoup@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com>
Co-authored-by: Lee Nau <lee.nau@gmail.com>
Co-authored-by: HMING <126185151+Hearum@users.noreply.github.com>
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: Hanming Lu <hanminglu@meta.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
|
2026-08-04 13:22:49 -07:00 |
|
 Oguz UlgenandCheng Wan
|
c113ead98a
|
Bump helion version to 1.4 (#32562)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-08-03 19:10:31 -07:00 |
|
Xiaoyu Zhang
|
c32c4ef79c
|
[Kernel] Move sgl-kernel under sglang.kernels.aot (#32648)
|
2026-07-29 17:25:00 +08:00 |
|
 Cheng WanandShu Wang
|
659d349b61
|
[core/loader] Add presharded load format (#24256)
Co-authored-by: Shu Wang <shuwanguc@google.com>
|
2026-07-25 13:03:39 -07:00 |
|
 
|
3079157175
|
Fix PyPI release: drop the git-only sgl-eval dep from packaged metadata (#32354)
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-24 20:00:57 -04:00 |
|
 
|
962c076934
|
Decode input_audio media containers with PyAV & Update memory profiler (#31832)
Signed-off-by: Shiyan Deng <dsy842974287@meta.com>
Signed-off-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
|
2026-07-24 16:18:45 -07:00 |
|
Baizhou Zhang
|
f15b43242b
|
Bump sgl-deep-gemm to 0.1.5 (#32345)
|
2026-07-24 14:03:38 -07:00 |
|
Rain Jiang
|
7fe82dd02e
|
create rust workspace (#32014)
|
2026-07-23 12:02:41 -07:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
99f636a86f
|
[Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 08:35:09 +08:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
0c29c8fece
|
Bump FlashInfer to 0.6.15.post1 (#31927)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-07-22 14:21:59 -07:00 |
|
Mohammad Miadh Angkad
|
3d82dacd58
|
Bump CuTe DSL to 4.6.0 (#31714)
|
2026-07-20 02:11:59 -07:00 |
|
       
|
02236fa38c
|
Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
|
2026-07-19 22:57:37 -07:00 |
|
Baizhou Zhang
|
304a529558
|
Revert "Bump FlashInfer to 0.6.15 and revert regressions" (#31625)
|
2026-07-17 16:46:33 -07:00 |
|
Lianmin Zheng
|
c95026aed3
|
Upgrade llguidance to 1.7.6 (#31484)
|
2026-07-17 16:31:44 -07:00 |
|
 sglang-botandsglang-bot
|
0ad0ff2e9e
|
chore: bump sglang-kernel version to 0.4.5 (#31618)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-17 15:58:36 -07:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
d67aa05697
|
Bump FlashInfer to 0.6.15 and revert regressions (#31502)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-07-17 00:50:12 -07:00 |
|
 Khoa PhamandClaude Fable 5
|
dc60f65661
|
chore: bump tokenspeed_mla to 0.1.8 (#31385)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 22:37:54 -07:00 |
|
     
|
423b8485fb
|
[Quantization] add humming quantization kernel (#23754)
Co-authored-by: guzekai01 <zekai01@antgroup.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@gmail.com>
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-14 08:42:56 +08:00 |
|
Mick
|
649ce5dd3d
|
model: support Pi0.5 (#30633)
|
2026-07-11 07:50:58 +08:00 |
|
Baizhou Zhang
|
3f1694f5e0
|
Update sgl-deep-gemm to 0.1.4.post1 (#30697)
|
2026-07-10 13:30:43 -07:00 |
|
Xinyuan Tong
|
b76dd0be69
|
Fix Mistral GSM8K chat eval (#27757)
|
2026-07-09 21:08:48 -07:00 |
|
 
|
2c6cd1ef41
|
[Dep] Upgrade flashinfer to 0.6.14 (#29910)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-07-09 17:52:33 -07:00 |
|
 Chetan Kumar VermaandMa Mingfei
|
b3ab56545b
|
Add Accuracy Benchmark for OCR models (#25364)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-06 16:18:43 +08:00 |
|
Baizhou Zhang
|
c312cdd3a7
|
Upgrading tvm-ffi/sgl-deep-gemm/tilelang (#29554)
|
2026-07-01 12:32:18 -07:00 |
|
Jialin Ouyang
|
40594bd381
|
[passthrough] engine: zstd request-body decompression + header overrides (#29684)
|
2026-06-30 23:48:56 -07:00 |
|
Mohammad Miadh Angkad
|
bae78a44da
|
[Deps] Bump transformers to 5.12.1 (#29393)
|
2026-06-30 00:54:01 -07:00 |
|
  
|
cfc0a0e0e0
|
Add Intel Quantization Support in SGLang (#18139)
Signed-off-by: Mengni Wang <mengni.wang@intel.com>
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Weiwei <weiwei1.zhang@intel.com>
|
2026-06-26 09:54:35 +08:00 |
|
 Khoa PhamandCursor
|
0642cd5020
|
(chore): bump tokenspeed_mla to 0.1.7 (#28759)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-21 21:25:37 -07:00 |
|
Mick
|
a51d56d948
|
CI: Pin flash-attn-4 for diffusion CI consistency (#28838)
|
2026-06-21 21:13:41 +08:00 |
|
Lianmin Zheng
|
8a3d6c3403
|
Sort pyproject dependency lists (#28811)
|
2026-06-20 17:24:57 -07:00 |
|
 sglang-botandsglang-bot
|
b1d18d562b
|
chore: bump sglang-kernel version to 0.4.4 (#28572)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-06-17 16:13:39 -07:00 |
|
Baizhou Zhang
|
27291118b9
|
Upgrade sgl-deep-gemm to 0.1.3 (#28402)
|
2026-06-16 23:06:36 -07:00 |
|
 feliang-gitandxutizhou
|
92b42c8d8a
|
LPLB: linear-programming load balancer for MoE expert parallelism (#24515)
Co-authored-by: xutizhou <xutingz@nvidia.com>
|
2026-06-16 10:19:42 -07:00 |
|
Xinyu Zhang
|
1800d7caa6
|
Bump ray minimum version to 2.55.1 (#27724)
|
2026-06-12 20:49:11 -07:00 |
|
Khoa Pham
|
a0c6e0b3a4
|
chore: bump tokenspeed_mla 0.1.1 -> 0.1.6 (#28116)
|
2026-06-12 19:57:36 -07:00 |
|
Mohammad Miadh Angkad
|
7f706f4cfb
|
[Deps] Bump FI to 0.6.12 and cutedsl to 4.5.2 (#26854)
|
2026-06-03 12:09:18 -07:00 |
|
Baizhou Zhang
|
68caf49154
|
Update sgl-deep-gemm to 0.1.2 (#26993)
|
2026-06-02 13:51:54 -07:00 |
|
Mick
|
64a1dec8b6
|
[diffusion] feat: add realtime webui super resolution controls (#27026)
|
2026-06-02 20:29:21 +08:00 |
|
Mick
|
3b26644bc4
|
[diffusion] misc: add realtime-webui (#26959)
|
2026-06-02 14:13:02 +08:00 |
|
Mick
|
2fc548f250
|
[diffusion] model: support lingot-world (#26954)
|
2026-06-02 13:52:49 +08:00 |
|