Commit Graph
498 Commits
Author SHA1 Message Date
Mick 22faf9fef8 embedding: centralize capabilities and complete OpenAI compatibility (#32481) 2026-07-30 10:28:52 +08:00
Mick 2aa86e9130 [diffusion] docs: add diffusion cookbook model tags (#32836) 2026-07-30 10:03:55 +08:00
amote-i 20d5b91e5b [NPU] [DOC] update feature name to follow the code changement (#32749) 2026-07-30 09:54:08 +08:00
zijiexiaandClaude Opus 5 5efbb18a6f [docs] Rotate popular models on the landing pages, lead the Cookbook nav with Kimi (#32835)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 17:46:50 -07:00
zijiexiaandClaude Opus 5 3c9efaf3e1 [docs] Kimi-K3: widen the H200 High-Throughput recipe to 4x8 TP32/EP32 (#32834)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 17:18:56 -07:00
Liangsheng Yin e5c46ff07d [Fix] Route asymmetric-KV models to fa4 on SM100 and pin MiMoV2 FP8 MoE to flashinfer_trtllm (#32818) 2026-07-29 16:37:43 -07:00
Ke Bao f73d6f2789 Update Inkling cookbook install command (#32799) 2026-07-30 02:23:23 +08:00
YAMY fddfc1fb5e [GDN] Support FlashInfer GDN prefill with extra-buffer radix cache (#29735) 2026-07-30 00:47:35 +08:00
Xiaoyu Zhang c32c4ef79c [Kernel] Move sgl-kernel under sglang.kernels.aot (#32648) 2026-07-29 17:25:00 +08:00
Andrew KuksaandANDREW_K 0caf0fc01d [Diffusion][Docs] Ascend A2, A3 add basic usage and benchmark results in diffusion cookbook (#30614)
Co-authored-by: ANDREW_K <andrewsha3@DESKTOP-KNDINTT.localdomain>
2026-07-29 11:50:15 +03:00
f05c92fb6d ✨ [llm][npu][quant] Add W8A8 MXFP8 quantization for Qwen3 MoE on Ascend NPU (#30768)
Co-authored-by: Артем Савкин <58187114+OrangeRedeng@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-29 10:39:36 +03:00
Mick 9a03bebf13 docs(kimi-k3): clarify VLM compatibility (#32661) 2026-07-29 11:53:04 +08:00
amote-i cb12a1547b [NPU] [DOC] update supported features on ascend npu (#32647) 2026-07-29 09:05:21 +08:00
YAMYandLee Nau 86ee545388 docs(cookbook): update Kimi-K3 GB200 recipes from measured 4x4 runs (#32592)
Co-authored-by: Lee Nau <lnau@nvidia.com>
2026-07-28 16:26:40 -07:00
Liangsheng Yin 85618cc798 docs: update sglang cookbook (#32654) 2026-07-28 15:33:24 -07:00
Qiaolin YuandZijie Xia 9c0dbf508f [cookbook] add inkling dspark command (#32465)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-07-28 14:33:38 -07:00
YAMY dd67452b4f [Cleanup] Move mamba-max-states-per-path validation into _handle_mamba_backend (#32502) 2026-07-28 14:21:21 -07:00
Mick 161fffedfc docs: clarify diffusion stage reuse guidance (#32639) 2026-07-28 19:52:52 +08:00
zijiexiaandClaude Opus 5 b79388f338 docs(cookbook): mark every Kimi-K3 cell in-progress; land Playground "Switch base" on the configurator (#32586)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 18:59:57 -07:00
sglang-botandsglang-bot edc0e5489f docs: sync LMSYS SGLang blog cards (#32590)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-27 18:00:29 -07:00
Yuhao Yang 7cae831e41 Update mi35x ROCm image to k3-20260727 (#32559) 2026-07-27 10:52:41 -07:00
Mohammad Miadh AngkadandXinyuan Tong 3ebb7c2d07 docs: point Kimi-K3 references to public branch (#32547)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-07-27 16:23:43 +00:00
7dafacca49 docs(cookbook): add the Kimi-K3 serving cookbook (#32542)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: thomawan <thomawan@amd.com>
Co-authored-by: BBuf <1182563586@qq.com>
2026-07-27 08:37:30 -07:00
8d6549bc40 [Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
2026-07-27 21:31:04 +08:00
Mick 08af5aea57 optimize: optimize EmbeddingGemma prefill performance (#32383) 2026-07-27 17:34:29 +08:00
Jackey HuaandClaude Opus 5 9a0bd24bed model: serve bare Qwen3Model backbone natively as an embedding model (#32457)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:49:58 +08:00
Mohammad Miadh Angkad 9791fc7090 Add configurable FlashInfer autotune skips (#31389) 2026-07-25 11:17:58 -07:00
Mohammad Miadh Angkad 953c587adf [Docs] Add Qwen3.6 35B NVFP4 to cookbook (#31413) 2026-07-25 11:16:53 -07:00
SovietPowerandShangming Cai 6a046fad09 [PD] Prevent decode scheduler from blocking on ZMQ sends to a stalled prefill peer (#31144)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-07-25 12:53:41 +08:00
sglang-botandsglang-bot 1ef2b85ef7 chore: bump docs install version to 0.5.16 (#32347)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-24 17:19:39 -07:00
Zhihao Wang 0a212c6119 [RL] DSV4: add env to quantize SWA KV cache from bf16-rounded values (#31086)
Signed-off-by: zhihaow6 <zhihaow6@illinois.edu>
2026-07-24 15:52:00 -07:00
YAMY de816e1eb5 [Disagg][StagingBuffer][1/2] Robustness and failure handling (#31217) 2026-07-24 17:22:55 +08:00
Polisetty V R K Jyothendra VarmaandMa Mingfei 319055c191 [Intel GPU] Add XPU Platform support (#31949)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-07-24 12:49:29 +08:00
Jinyan Yi b98a577fbe Doc/update ascend quickstart (#32205) 2026-07-23 20:25:20 +08:00
Liangsheng Yin 9b853e6832 [Scheduler] Enable decode retraction ordering under speculative decoding (#32023) 2026-07-23 00:42:56 -07:00
sglang-botandsglang-bot 9bda6fdb9a docs: sync LMSYS SGLang blog cards (#32127)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-23 00:57:01 +00:00
Xiaoyu ZhangandClaude Opus 4.8 99f636a86f [Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 08:35:09 +08:00
Brayden Zhong e7511141ea Support CuteDSL GEMM BF16 on SM100 on by default when allowed by heuristic (#30567) 2026-07-22 14:13:12 -07:00
0a6d1930c3 [Attention Backend] Add HPC-Ops attention backend (#30540)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
2026-07-22 22:06:22 +08:00
axx-ty911 7dcf3f7cbf [Doc] Rename A5 product name (#31998) 2026-07-22 15:09:05 +08:00
4a55fdba0b docs(cookbook): re-benchmark DeepSeek-V4 on sglang 0.5.15 (#31363)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-07-21 22:55:36 +00:00
9a6c96083f [Cookbook] Add Laguna-S-2.1 (#31918)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-21 17:01:50 +00:00
dcd9014f15 [AMD][MXFP4] Reland "Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28291)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-07-21 02:50:23 -07:00
Rahul Vijayaraghavan fa0ced195e [XPU] Enable breakable prefill CUDA graph on XPU (#30273) 2026-07-21 09:09:40 +08:00
zijiexiaandClaude Opus 4.8 0a2d3ca071 [cookbook] Inkling: add measured accuracy numbers to benchmark cards (#31823)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 22:43:35 +00:00
Douglas YangandClaude Opus 4.8 e149cdb337 docs(cookbook): revert MiniMax-M3 to dev image (model not yet in a release) (#31819)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 12:16:47 -07:00
Ke Bao 7fc545b649 Align reasoning_effort schema across chat, tokenize, and responses (#31784) 2026-07-20 23:22:43 +08:00
amote-i 97e0647bdc [NPU] [DOC] Fix issues about npu docs found by aidd (#31302) 2026-07-20 16:31:27 +08:00
jacky.cheng 17fdd8487f [AMD] Update qwen3.5 cookbook (#31737) 2026-07-20 14:14:14 +08:00
b8ec544946 [DSA] Integrate Q8KV8 FP8 Sparse MLA Prefill into the DSA Backend (DeepSeek-V3.2) (#30514)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-07-19 11:58:16 +08:00