Commit Graph
590 Commits
Author SHA1 Message Date
Trevor Morris a16872767f Update CUDA 13.4 image to flashinfer 0.6.18rc10, cutedsl 4.8. Fix sgl- wheel unpinning (#36929) 2026-08-28 17:52:15 -07:00
Bingxu Chen 9579bff860 [AMD] Add ROCm 10 (gfx942 / gfx950) release images (#36434) 2026-08-28 08:03:39 -07:00
Trevor Morris ce6e1f46b4 Limit concurrent build jobs for cu134 container (#36756) 2026-08-27 20:24:35 -07:00
cbd0271574 [diffusion] model: add cosmos3 transfer capability (#34747)
Co-authored-by: Kedi Wu <kediw@nvidia.com>
Co-authored-by: Kedi Wu <31940276+kediwu0331@users.noreply.github.com>
2026-08-28 09:37:22 +08:00
Trevor Morris 3ce243da3f [NVIDIA] Add CUDA 13.4 container for initial Rubin support (#36233) 2026-08-26 15:02:16 -07:00
Shangming Cai c8b56b1f44 chore: bump mooncake version to 0.3.13 (#36493) 2026-08-26 22:21:15 +08:00
Артем Савкинandronnie_zheng 61b67316d8 [NPU] [Diffusion] Support MiniMax H3 on Ascend NPU's (#33569)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-08-25 09:44:14 +03:00
514b997e6c Register CPU CI for 17 e2e tests and partition xeon base-c suite (#35227)
Co-authored-by: Zhang, Mingxu <mingxu.zhang@intel.com>
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-24 10:49:38 +08:00
b98d472158 [NPU] [Diffusion] Fix critical Ascend NPU Diffusion regression/bugs & restore 2-NPU CI testcase (#34855)
Co-authored-by: Elizaveta Martirosian <elizabet3000@mail.ru>
Co-authored-by: Arseniy Mironov <98156294+Napkin-AI@users.noreply.github.com>
Co-authored-by: Alexandr <117110413+Allor-maker@users.noreply.github.com>
Co-authored-by: P_Alex_Tr <aleksandr.smyshlaev@yandex.ru>
2026-08-22 18:32:20 +03:00
Bingxu Chen 6a12583679 [AMD] Update ROCm AITER pin to c16d44b (#35810) 2026-08-21 00:02:17 -07:00
Bingxu Chen 78c964d9d7 [AMD] Retry transient network failures in ROCm Dockerfile curl fetches (#35654) 2026-08-20 21:54:24 -07:00
Baizhou Zhang 92eeed41d7 [Docker] Defer CUDA 13 NCCL override until after dependency resolution (#35756) 2026-08-20 14:48:07 -07:00
Shangming Cai 628674a5c7 Remove unused MOONCAKE_COMPILE_ARG argument from Dockerfile (#35649) 2026-08-20 14:19:52 +08:00
360d10d6bc [Feature] Add process-local in-memory KV indexer and Router integration (#33370)
Co-authored-by: Wu, Yutong <yutong.wu@amd.com>
Co-authored-by: TianDi101 <ditian12@amd.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-08-20 10:45:35 +08:00
datdo-msft f65961844a docker: fix CUDA-13 build — rename NCCL_VERSION ARG to avoid base image ENV collision (#35587) 2026-08-19 18:41:53 -07:00
chuyehandChen c7478228dd [AMD] [Docker] Upgrade Python 3.12 + torch 2.11 + triton 3.7 in ROCm 7.2.4 (#30984)
Co-authored-by: Chen <bingxche@amd.com>
2026-08-19 18:18:31 -07:00
Augusto Yao 3a8f522f65 install sglang in virtual env instead of system path (#30612) 2026-08-18 16:36:50 -07:00
Baizhou Zhang cfc6dfb364 Apply latest DeepEP branch (#34923) 2026-08-18 16:24:50 -07:00
ashwini rathi 5eb117f6ca [XPU] Fix decode graph runner is_current_stream_capturing on non-CUDA devices (#35050) 2026-08-18 14:28:13 +08:00
744740dbea [XPU] upgrade sglang xpu backend to PyTorch 2.13 (#31751)
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-17 18:29:15 +08:00
MingxuZh eafbe2cb6f [CI] Install sgl-eval in xeon (CPU) Docker image (#34818) 2026-08-17 14:00:44 +08:00
Lianmin Zheng 67e12131df Build Rust extensions on demand in source checkouts (#34994) 2026-08-16 14:58:06 -07:00
Baizhou Zhang 8d44091326 Update sgl-deep-ep release workflow for DeepEP v2 (#34914) 2026-08-15 02:08:28 -07:00
Brayden ZhongandBrayden Zhong 5e65dd01a7 Remove the torchao integration (--torchao-config) (#34304)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-14 21:49:11 +08:00
Mohammad Miadh Angkad 00e57d74f0 Bump FlashInfer to 0.6.17 and remove Kimi K3 workarounds (#33997) 2026-08-12 02:17:26 -07:00
Bingxu ChenandYC Yen-Ching Tseng 6f3fe13a9c [AMD] Install AITER's pinned Triton wheel in the ROCm 7.2 image (#34364)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-08-11 16:36:45 +08:00
Mick c20e99bd22 fix(vlm): preserve Kimi-K3 GPU JPEG accuracy (#34163) 2026-08-10 09:42:52 +08:00
Baizhou Zhang c9444deef4 Docker: install DeepEP from release wheels (#34041) 2026-08-07 23:51:09 -07:00
Baizhou Zhang 8fc66d1a62 docker: pin Kimi images to SGLang commit (#34030) 2026-08-07 14:49:26 -07:00
Baizhou Zhang 8a22b8305d docker: add Kimi K3 artifacts and build hpc-ops with C++20 (#33956) 2026-08-07 13:50:44 -07:00
Baizhou Zhang e0af47b03e Fix sgl-deep-ep builder dependencies (#33866) 2026-08-06 16:30:02 -07:00
Mohammad Miadh AngkadandBrayden Zhong 434e646282 [Deps] Upgrade CUDA PyTorch stack to 2.13 (#28836)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-08-06 12:08:44 -07:00
Baizhou Zhang ba9074035a Build and release sgl-deep-ep wheels (#33498) 2026-08-06 02:07:36 -07:00
Baizhou Zhang 25035bff8d Use main branch in Kimi Dockerfiles (#33760) 2026-08-05 15:33:14 -07:00
Zhaoyi Li b327d76682 [AMD] Bump mori to latest in sglang (#33462) 2026-08-04 15:12:40 -07:00
+26 abddb1c7e9 [Kimi] Support kimi-k3 (#32541)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai>
Co-authored-by: Zijie Xia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: zhangxiaohao <1024393531@qq.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: RolaoDenthu <xinyisong0111@gmail.com>
Co-authored-by: pigeonsoup <32922982+pigeonsoup@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com>
Co-authored-by: Lee Nau <lee.nau@gmail.com>
Co-authored-by: HMING <126185151+Hearum@users.noreply.github.com>
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: Hanming Lu <hanminglu@meta.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
2026-08-04 13:22:49 -07:00
ashwini rathiandMa Mingfei 53804d609c [CI][XPU] Stabilize XPU CI: pin UMD/IGC, retry infra flakes, right-size EAGLE3 (#32438)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-04 16:28:52 +08:00
Mohammad Miadh Angkad a1344fad4e Replace Kimi K3 DeepGEMM patch with 0.1.5.post1 (#33143) 2026-07-31 22:17:46 -07:00
Bingxu ChenandCursor Agent 48c1b37a33 [AMD] Update ROCm AITER pin to d9e5ef7 (#32939)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-07-30 22:05:03 +08:00
Bingxu Chen 3d6e1e6f81 [AMD] Revert ROCm AITER pin to 9127c94 (#32879) 2026-07-30 11:53:49 +08:00
Baizhou Zhang d12ea3e9ba docker: add Kimi K3 images (#32760) 2026-07-29 03:02:28 -07:00
Xiaoyu Zhang c32c4ef79c [Kernel] Move sgl-kernel under sglang.kernels.aot (#32648) 2026-07-29 17:25:00 +08:00
8d6549bc40 [Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
2026-07-27 21:31:04 +08:00
Shangming Cai bdd0698541 chore: bump mooncake version to 0.3.12.post1 (#32302) 2026-07-25 14:12:46 +08:00
Baizhou Zhang f15b43242b Bump sgl-deep-gemm to 0.1.5 (#32345) 2026-07-24 14:03:38 -07:00
Baizhou Zhang ed26a111b2 Revert "docker: install dynamo nightly in the dev image for rapid iteration/testing" (#32257) 2026-07-23 14:51:43 -07:00
Mohammad Miadh AngkadandBrayden Zhong 0c29c8fece Bump FlashInfer to 0.6.15.post1 (#31927)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-07-22 14:21:59 -07:00
monkeyLoveding d1c2a1de08 [NPU] memfabric-zbal update (#31777) 2026-07-20 22:01:28 +08:00
02236fa38c Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
2026-07-19 22:57:37 -07:00
Baizhou Zhang 304a529558 Revert "Bump FlashInfer to 0.6.15 and revert regressions" (#31625) 2026-07-17 16:46:33 -07:00