Commit Graph
351 Commits
Author SHA1 Message Date
sglang-botandsglang-bot 081a01c37d docs: sync LMSYS SGLang blog cards (#29307)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-30 20:48:10 +00:00
97fc4dfd73 [Doc]Checking and modifying Markdown formatting issues and link validity (#28586)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-30 20:27:38 +03:00
Yihao WangandClaude Opus 4.8 a531d81c19 [diffusion][cache-dit] support Krea-2 + run-driven has_separate_cfg (#29688)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 08:22:05 -07:00
Trevor Morris b8c25bfaa7 [NVIDIA] Support flashinfer a2a with flashinfer_trtllm_routed moe (#22394) 2026-06-29 16:23:58 -07:00
Cheng Wanandlch1475369 fc96edd297 feat(mem_cache): page-major (layer-major within a page) KV/state layout (#29533)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-06-29 14:49:54 -07:00
zijiexiaandClaude Opus 4.8 5106b42cbd [cookbook] GLM-5.2 NVFP4 B300: TP8 recipe + 3 strategies (#29557)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:42:19 -07:00
zijiexiaandClaude Opus 4.8 11b7ed7c9e [Docs] Add --prerelease=allow so uv installs the latest sglang (#29676)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:13:34 -07:00
zijiexiaandClaude Opus 4.8 74a197af9d docs: add B200 NVFP4 recipes + benchmarks to GLM-5.2 cookbook (#29674)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:06:30 -07:00
Jyothirmai KottuandXinyuan Tong 473a278dd1 model: support nvidia/LocateAnything-3B (#28958)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-30 00:16:42 +08:00
amote-i 489017b3d6 [NPU] [DOC] Update deterministic inference feature support status to A2, A3 (#29632) 2026-06-29 19:17:06 +08:00
danielafrimiandDaniel Afrimi a2b5ce2ed1 Add stochastic rounding for FP16 Mamba SSM cache (#26929)
Signed-off-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
2026-06-29 01:47:09 -07:00
Xinyuan Tong 38d4ffcd86 [cookbook] drop redundant serve flags (GLM-5.2) + fix M3 page-size note (#28731) 2026-06-29 13:34:13 +08:00
jianzhao-xu 2260e612f6 [NPU] update best practicce docs from testcase (#29492) 2026-06-29 11:30:27 +08:00
Liangsheng Yin 909123ddb8 [misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591) 2026-06-28 18:38:28 -07:00
Mick 643e1cc779 fix: fix prefill-aware SWA floor tracking (#29520) 2026-06-28 14:18:28 +08:00
Mick cfd911ad6e docs: refine diffusion cookbook overview (#29507) 2026-06-28 00:10:21 +08:00
Aditya KamatandMick 1589603114 model: support baidu unlimited-ocr (#29186)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-27 23:36:19 +08:00
b030b1a5f3 hisparse: support NIXL DRAM KV destinations for HiSparse (#27563)
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-27 22:32:29 +08:00
jonah-bermanandXinyuan Tong 2f34dbe372 Add native Exa-backed web_search support (#29342)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-27 14:41:09 +01:00
amote-i a3c5e286f6 [NPU] [DOC] Fix and update Ascend NPU docs (#29501) 2026-06-27 18:23:11 +08:00
zijiexiaandClaude Opus 4.8 e0c0c0a45c [Cookbook] GLM-5.2: tune GB300 NVFP4 recipes + fill benchmarks (#29486)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 01:01:28 -07:00
Baizhou Zhang 12f76d115c Update GLM-5.2 B300 and GB300 NVFP4 cookbook settings (#29466) 2026-06-26 16:15:33 -07:00
Yaochen Hanandronnie_zheng c98d31143d update quantization code owner and document quantization contributions (#26784)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-26 19:40:33 +03:00
zijiexiaandClaude Opus 4.8 c21c2d9421 [Docs] Cookbook: match playground docker image resolution to deployment (hw|quant) (#29400)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 00:56:34 -07:00
aeb4e98108 [diffusion] model: support JoyEcho multi-shot A/V generation support (#27420)
Co-authored-by: niehen6174 <niehen6174@users.noreply.github.com>
Co-authored-by: 1639206518@qq.com <niehen6174>
2026-06-26 15:46:31 +08:00
zijiexiaandClaude Opus 4.8 dd56a9f069 [Docs] Add NVFP4 quantization to GLM-5.2 cookbook (#29380)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 00:40:13 -07:00
jianzhao-xu cc294829aa [NPU] fix best practicce docs (#29303) 2026-06-26 11:17:50 +08:00
Raiden MakotoandRaiden-Makoto 7f376644e0 [AMD] [GLM5] Mark EAGLE verified on MI300X/MI325X (gfx942) in GLM-5.1 cookbook (#29313)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-26 10:49:58 +08:00
cfc0a0e0e0 Add Intel Quantization Support in SGLang (#18139)
Signed-off-by: Mengni Wang <mengni.wang@intel.com>
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Weiwei <weiwei1.zhang@intel.com>
2026-06-26 09:54:35 +08:00
amote-i 10ff3c1dcb [NPU] [DOC] Add environment prerequisites to model tutorials (#29293) 2026-06-26 09:51:06 +08:00
Yi Zhong f83cbc2516 Add LFM2.5-230M to the LFM2.5 cookbook (#29321) 2026-06-26 08:27:57 +08:00
Xiaoyu Zhang 4d06d4c97f Sync Gemma4 hardware table with Blackwell recipes (#29266) 2026-06-25 22:57:33 +08:00
52c32035eb [diffusion] Add Qwen-Image ModelOpt NVFP4 support (#28928)
Co-authored-by: jingyu-ml <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 22:56:09 +08:00
Mick 890b38c211 [diffusion] doc: fix diffusion docs and cookbook drift (#29302) 2026-06-25 21:13:00 +08:00
Raiden MakotoandRaiden-Makoto 0075c8f02b [AMD] [GLM5] GLM-5.1 MXFP4 (MI355X) + enable EAGLE for gfx950 in cookbook (#29194)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-25 03:54:48 -07:00
Mohammad Miadh Angkad e4976683f4 [Docs] Fix broken links in cookbook (#29261) 2026-06-25 02:31:55 -07:00
Xiaoyu Zhang 7c9804ef21 Add MiMo V2.5 Blackwell vision FA4 recipe (#29253) 2026-06-25 13:47:32 +08:00
Xiaoyu Zhang efbe67d237 Tune Gemma4 26B-A4B B200 memory recipe (#29252) 2026-06-25 11:31:25 +08:00
Brayden ZhongandBrayden Zhong f82addd4a8 Support online MXFP8 quantization for ungated MoE (#27939)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-24 16:58:48 -07:00
Brayden ZhongandBrayden Zhong 2c697daf5f [Cookbook] Nemotron3-Ultra: align MTP draft depth with NVIDIA reference (num_steps 5) (#29200)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-24 19:12:34 +00:00
Anusha Pant e97cc339e3 Add DeepSeek V4 Flash demo notebook (#28952) 2026-06-24 18:42:33 +00:00
Yihao Wang dd4caf9459 [Docs] fix SGLang-diffusion installation links (#29139) 2026-06-24 09:46:06 -07:00
amote-i 73d976e375 [NPU] [DOC] Fix TOC of Ascend NPU Docs (#29129) 2026-06-24 16:47:10 +08:00
zijiexiaandClaude Opus 4.8 dd2d919e21 [Docs] Fix mem-fraction-static default and document how it is computed (#29135)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 08:05:06 +00:00
Trevor Morris f74a1722e6 [NVIDIA] Support TF32 matmul to improve MiniMax gate gemm performance (#22744) 2026-06-23 14:54:54 -07:00
Yihao Wang 83d32fbc2f remove lora section (#29056) 2026-06-23 16:26:51 +00:00
Yihao Wang b5e0965b07 [Diffusion][Cookbook] Add Krea-2 cookbook (#29051) 2026-06-23 08:02:08 -07:00
zijiexiaandClaude Opus 4.8 52a90c9a36 docs(minimax-m3): use published AMD ROCm images (#28777)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 07:48:16 +01:00
Thomas Wang 7e6587c94a [AMD] Update v4 cookbook to clean env vars (#28981) 2026-06-22 22:16:29 -07:00
amote-i 84338df6f0 [NPU] [DOC] Update contribution guide of Ascend NPU (#28909) 2026-06-23 10:00:12 +08:00