Commit Graph
1402 Commits
Author SHA1 Message Date
Mick 46d84f4b48 feat(cli): add extensible serve backend plugins (#34753) 2026-08-14 13:57:59 +08:00
zijiexiaandClaude Opus 5 463981922c [Cookbook] Add DeepSeek-V4-Pro-0813 (Pro Official) serving recipes (#34809)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:53:41 -07:00
Yuhao YangandClaude 6ad3f2d8fd docs: link dots3.note checkpoints, add H100 cells (#34797)
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-14 02:42:27 +00:00
Ziang Li 9d34c2809f [FlashInfer v0.6.16] Support FlashInfer CuTe DSL NVFP4 MoE quantization (#28354) 2026-08-13 17:33:46 -07:00
zijiexia abdef3c38e [Qwen] Update Docker image tag for MI300X to v0.5.17-rocm700-mi30x-20… (#34770) 2026-08-13 15:59:42 -07:00
Khoa PhamandCursor 652a2709d1 [Docs] Add decode context parallelism to advanced features (#34654)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-13 13:51:57 -07:00
c1142677a8 [do not merge] add new cookbooks (#34658)
Co-authored-by: Jianfei Wang <jianfei.wangg@outlook.com>
Co-authored-by: Jianfei Wang <905787410@qq.com>
2026-08-13 13:35:15 +00:00
fad376d3ee [CPU][QUANT] add amx cpu support for auto-round (#29593)
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Signed-off-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: sys-lpot-val <sys_lpot_val@intel.com>
Co-authored-by: Weiwei Zhang <WeiweiZhang1@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-13 15:50:59 +08:00
Carrie ChenandBrayden Zhong 6a5a9eccaa add flashinfer cute-dsl backend for mxfp8 gemm (#34042)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-13 08:50:01 +08:00
Lukas Humbel 40eaf34428 fix: make automatic NUMA binding configurable (#30394) 2026-08-12 17:40:28 -07:00
Jimmy Shong 198b7e9240 [Docs] Use Meta's canonical Muse Glimmer GGUF filename (#34626) 2026-08-12 16:07:55 -07:00
YAMY e6250c7c70 docs: update Qwen3.8 disaggregated serving configs (#34601) 2026-08-12 11:02:35 -07:00
Xinyuan Tong d21eefc94f [Docs] Rename Qwen3.8-Max-DSpark to Qwen3.8-2.4T-A95B-DSpark (#34590) 2026-08-12 15:25:48 +00:00
Yichi Zhang f28bc5a6de docs(cookbook): add BF16 recipes to Nemotron 3.5 Lightning (#34573) 2026-08-12 08:15:34 -07:00
zijiexiaandClaude Opus 5 8e7c07fae7 [Docs] Add Qwen3.8 cookbook (#34587)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 08:07:46 -07:00
ajith-sirra-amdandgiovanniguastiamd b5d1453ed2 [AMD] GLM 5.2 MXFP4 SGLANG COOKBOOK (#34379)
Signed-off-by: Sirra <asirra@amd.com>
Co-authored-by: giovanniguastiamd <giovanni.guasti@amd.com>
2026-08-12 03:39:25 -07:00
Mick 644d55ebfa [diffusion] feat: support native and peft minimax h3 loras (#34359) 2026-08-12 17:52:40 +08:00
Mohammad Miadh Angkad 00e57d74f0 Bump FlashInfer to 0.6.17 and remove Kimi K3 workarounds (#33997) 2026-08-12 02:17:26 -07:00
Yuzhen Zhou 2d76d537e5 feat: support deterministic FA4 for GLM-4.7-Flash (#33945) 2026-08-12 16:57:34 +08:00
Mick a9a355774a [diffusion] feat: support dynamically cpu offload components (#34391) 2026-08-12 11:38:27 +08:00
Mick 2be9773a21 [diffusion] doc: update cosmos3 edge and distilled cookbook (#34497) 2026-08-12 10:51:11 +08:00
Xiaoyu Zhang a53d3636ce [diffusion][model] Add native SANA-Video T2V support (#32921) 2026-08-12 10:07:24 +08:00
Faradawn Yang 857910bd35 docs(semianalysis): Update Qwen3.5 B200 NVFP4 MTP config (#34357) 2026-08-11 16:12:17 -07:00
gongwei1027 2c07ca5e8d [Fix] Allow flashinfer_sparse_mla DSA backend for HiSparse on SM120 FP8 KV (#33075) 2026-08-11 15:05:04 -07:00
Douglas YangandClaude Opus 5 59450c4f18 docs(cookbook): Kimi-K3 — drop --enable-symm-mem from the GB cells (#34444)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 19:23:41 +00:00
Xinyuan Tong d5d41d07ed [Docs] Add Ling-3.0-tiny INT4 recipes (#34395) 2026-08-12 03:16:03 +08:00
Mick 8267d76c2c [VLM] replace deprecated image processor use_fast (#34175) 2026-08-12 00:14:07 +08:00
Jinyan Yi f148eb6e6e Add Hunyuan3 On Ascend Doc (#30223) 2026-08-11 21:20:30 +08:00
Faradawn YangandRyan Stewart 3add7e19ff Add NVIDIA Nemotron 3.5 Lightning cookbook (#33481)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
Signed-off-by: Ryan Stewart <rystewart@nvidia.com>
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-11 06:00:57 -07:00
Zhiqiang XieandTingwei Huang 5469faec45 HiSparse: shared-index (IndexShare) plan-then-IO swap-in prefetch (#34329)
Co-authored-by: Tingwei Huang <huangtingwei9988@gmail.com>
2026-08-11 01:58:28 -07:00
Xinyuan Tong 1c06c160f9 [Docs] Add Ling-3.0-flash INT4 and MXFP4 recipes (#34363) 2026-08-11 00:35:27 -07:00
Mick aeab1de1de [diffusion] optimization: support cuda graph for Pi-0.5 prefix encoding (#34256) 2026-08-11 09:14:53 +08:00
Baizhou Zhang a92bbf2f24 docs: remove DSV4 low-latency chunked prefill size (#34333) 2026-08-10 17:23:43 -07:00
Brayden ZhongandBrayden Zhong b86c90215f [Docs] Muse Glimmer cookbook: drop --speculative-dflash-block-size 5 (#34323)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-08-10 15:46:34 -07:00
e54c153ba6 Add Intern-S2-Mobius cookbook (#33820)
Co-authored-by: Justin Tong <justintong0323@outlook.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-10 13:13:35 -07:00
d07ac32d05 [diffusion] feat: support --served-model-name in sglang serve (#34228)
Co-authored-by: TobyMint <tobymint@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-10 22:27:09 +08:00
Xinyuan Tong 77c90e7e54 Cookbook: add Ling-3.0-tiny (#34283) 2026-08-10 21:25:17 +08:00
Mick 8ba9385097 [diffusion] chore: optimize model weight loading (#34064) 2026-08-10 20:07:48 +08:00
Brayden ZhongandBrayden Zhong d95e824a49 Muse Glimmer Cookbook: install from the PR branch (#34281)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-08-10 11:39:55 +00:00
Brayden ZhongandBrayden Zhong d96b1533ea Muse Glimmer Cookbook: install from the PR branch (#34278)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-08-10 03:49:02 -07:00
a6c34df044 Muse Glimmer Cookbook (#34271)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: Jimmy Shong <jimmysh341@gmail.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-08-10 10:21:13 +00:00
Mick 955569a2dc [diffusion] feat: expose cosmos3 policies through the Action API (#34243) 2026-08-10 18:16:20 +08:00
Ke Bao 430f38ea25 Update dspark draft path in Inkling small cookbook (#34250) 2026-08-10 17:12:51 +08:00
Feng Su fb3d1419fd [tracing] sglang tracing v2: support exporting tracing data asynchronously (#30023) 2026-08-10 15:23:05 +08:00
2969ab3d41 [MLX] Window-bounded SWA KV storage and in-graph sampling (#34166)
Co-authored-by: Siming Deng <siming_deng_stat@163.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
Co-authored-by: Jiminator <Jiminator@users.noreply.github.com>
Co-authored-by: damahua <damahua@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 21:28:51 -07:00
Mick 169783d42f [diffusion] chore: make torch.compile opt-in for speed mode (#34173) 2026-08-10 10:22:16 +08:00
Liangsheng Yin 7c90840bad [CI] Key scheduled CUDA suites by runner_config instead of hand-written jobs (#34186) 2026-08-09 16:44:53 -07:00
WenhaoZhang 51470b376f [diffusion] feat: support sol-attn sparse attention backend for h3 (#33702) 2026-08-09 16:26:06 +08:00
548ff545c5 [diffusion] fix: guard sage attention sm90 bindings (#34107)
Co-authored-by: RunFMe <RunFMe@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-08-08 21:32:40 +08:00
Mick cf2d4fd679 docs: clarify K3 VLM feature transport (#34099) 2026-08-08 19:23:00 +08:00