Commit Graph
365 Commits
Author SHA1 Message Date
ming_wang d8a4f7a7aa add mimo-v2-flash model tutorial (#29932) 2026-07-03 11:02:24 +08:00
Peng Xingchen a6bc7fef90 glm5.2 on ascend doc (new version) (#29828) 2026-07-03 10:19:44 +08:00
amote-i 70a813493f [NPU] [DOC] add missing DEEP_NORMAL_MODE_USE_INT8_QUANT for w8a8+deepep scenarios (#29937) 2026-07-03 10:18:22 +08:00
sglang-bot 91d7645aca docs: sync LMSYS SGLang blog cards (#29990) 2026-07-03 01:01:32 +00:00
Jimmy Shong 85e71b7e13 [Doc] Cookbook Laguna-XS-2.1: add AIME25 accuracy (B300 + GB300) (#29974) 2026-07-02 13:15:01 -07:00
zijiexiaandClaude Opus 4.8 cba3801f52 docs: add PD disaggregation to GLM-5.2 cookbook playground (#29544)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 09:57:38 -07:00
Jimmy Shong 476c946543 [Doc] Cookbook: Laguna-XS-2.1 (DFlash low-latency + high-throughput) (#29884) 2026-07-02 20:05:33 +08:00
zijiexiaandClaude Opus 4.8 1c75243f5e docs: add Qwen3.6-27B-NVFP4 variant to cookbook (#29905)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 00:05:29 -07:00
Cao E 926140d789 [XPU] Enable XPU graph support (decode full-graph + prefill tc_piecewise) (#29053) 2026-07-02 13:24:35 +08:00
Zaili Wang cb06c4e6ce [CPU] Fix model failures on Xeon (#29497) 2026-07-02 13:20:18 +08:00
Cheng Wanandlch1475369 4a8e76805c feat(mem_cache): unified memory pool for hybrid Mamba / SWA models (#29678)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-07-01 13:21:59 -07:00
Baizhou Zhang 677a11bfa9 [Doc] Tiny update dsv4 doc (#29827) 2026-07-01 01:47:35 -07:00
qinsir5522 a7390b17f8 [NPU]Modify --lora-backend & --moe-runner-backend description. (#29793) 2026-07-01 11:15:46 +08:00
bb98629157 docs(cookbook): add AMD MI300X/MI325X/MI355X support for GLM-5.2 (#28471)
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-30 21:08:51 +00:00
sglang-botandsglang-bot 081a01c37d docs: sync LMSYS SGLang blog cards (#29307)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-30 20:48:10 +00:00
97fc4dfd73 [Doc]Checking and modifying Markdown formatting issues and link validity (#28586)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-30 20:27:38 +03:00
Yihao WangandClaude Opus 4.8 a531d81c19 [diffusion][cache-dit] support Krea-2 + run-driven has_separate_cfg (#29688)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 08:22:05 -07:00
Trevor Morris b8c25bfaa7 [NVIDIA] Support flashinfer a2a with flashinfer_trtllm_routed moe (#22394) 2026-06-29 16:23:58 -07:00
Cheng Wanandlch1475369 fc96edd297 feat(mem_cache): page-major (layer-major within a page) KV/state layout (#29533)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-06-29 14:49:54 -07:00
zijiexiaandClaude Opus 4.8 5106b42cbd [cookbook] GLM-5.2 NVFP4 B300: TP8 recipe + 3 strategies (#29557)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:42:19 -07:00
zijiexiaandClaude Opus 4.8 11b7ed7c9e [Docs] Add --prerelease=allow so uv installs the latest sglang (#29676)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:13:34 -07:00
zijiexiaandClaude Opus 4.8 74a197af9d docs: add B200 NVFP4 recipes + benchmarks to GLM-5.2 cookbook (#29674)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:06:30 -07:00
Jyothirmai KottuandXinyuan Tong 473a278dd1 model: support nvidia/LocateAnything-3B (#28958)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-30 00:16:42 +08:00
amote-i 489017b3d6 [NPU] [DOC] Update deterministic inference feature support status to A2, A3 (#29632) 2026-06-29 19:17:06 +08:00
danielafrimiandDaniel Afrimi a2b5ce2ed1 Add stochastic rounding for FP16 Mamba SSM cache (#26929)
Signed-off-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
2026-06-29 01:47:09 -07:00
Xinyuan Tong 38d4ffcd86 [cookbook] drop redundant serve flags (GLM-5.2) + fix M3 page-size note (#28731) 2026-06-29 13:34:13 +08:00
jianzhao-xu 2260e612f6 [NPU] update best practicce docs from testcase (#29492) 2026-06-29 11:30:27 +08:00
Liangsheng Yin 909123ddb8 [misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591) 2026-06-28 18:38:28 -07:00
Mick 643e1cc779 fix: fix prefill-aware SWA floor tracking (#29520) 2026-06-28 14:18:28 +08:00
Mick cfd911ad6e docs: refine diffusion cookbook overview (#29507) 2026-06-28 00:10:21 +08:00
Aditya KamatandMick 1589603114 model: support baidu unlimited-ocr (#29186)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-27 23:36:19 +08:00
b030b1a5f3 hisparse: support NIXL DRAM KV destinations for HiSparse (#27563)
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-27 22:32:29 +08:00
jonah-bermanandXinyuan Tong 2f34dbe372 Add native Exa-backed web_search support (#29342)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-27 14:41:09 +01:00
amote-i a3c5e286f6 [NPU] [DOC] Fix and update Ascend NPU docs (#29501) 2026-06-27 18:23:11 +08:00
zijiexiaandClaude Opus 4.8 e0c0c0a45c [Cookbook] GLM-5.2: tune GB300 NVFP4 recipes + fill benchmarks (#29486)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 01:01:28 -07:00
Baizhou Zhang 12f76d115c Update GLM-5.2 B300 and GB300 NVFP4 cookbook settings (#29466) 2026-06-26 16:15:33 -07:00
Yaochen Hanandronnie_zheng c98d31143d update quantization code owner and document quantization contributions (#26784)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-26 19:40:33 +03:00
zijiexiaandClaude Opus 4.8 c21c2d9421 [Docs] Cookbook: match playground docker image resolution to deployment (hw|quant) (#29400)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 00:56:34 -07:00
aeb4e98108 [diffusion] model: support JoyEcho multi-shot A/V generation support (#27420)
Co-authored-by: niehen6174 <niehen6174@users.noreply.github.com>
Co-authored-by: 1639206518@qq.com <niehen6174>
2026-06-26 15:46:31 +08:00
zijiexiaandClaude Opus 4.8 dd56a9f069 [Docs] Add NVFP4 quantization to GLM-5.2 cookbook (#29380)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 00:40:13 -07:00
jianzhao-xu cc294829aa [NPU] fix best practicce docs (#29303) 2026-06-26 11:17:50 +08:00
Raiden MakotoandRaiden-Makoto 7f376644e0 [AMD] [GLM5] Mark EAGLE verified on MI300X/MI325X (gfx942) in GLM-5.1 cookbook (#29313)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-26 10:49:58 +08:00
cfc0a0e0e0 Add Intel Quantization Support in SGLang (#18139)
Signed-off-by: Mengni Wang <mengni.wang@intel.com>
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Weiwei <weiwei1.zhang@intel.com>
2026-06-26 09:54:35 +08:00
amote-i 10ff3c1dcb [NPU] [DOC] Add environment prerequisites to model tutorials (#29293) 2026-06-26 09:51:06 +08:00
Yi Zhong f83cbc2516 Add LFM2.5-230M to the LFM2.5 cookbook (#29321) 2026-06-26 08:27:57 +08:00
Xiaoyu Zhang 4d06d4c97f Sync Gemma4 hardware table with Blackwell recipes (#29266) 2026-06-25 22:57:33 +08:00
52c32035eb [diffusion] Add Qwen-Image ModelOpt NVFP4 support (#28928)
Co-authored-by: jingyu-ml <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 22:56:09 +08:00
Mick 890b38c211 [diffusion] doc: fix diffusion docs and cookbook drift (#29302) 2026-06-25 21:13:00 +08:00
Raiden MakotoandRaiden-Makoto 0075c8f02b [AMD] [GLM5] GLM-5.1 MXFP4 (MI355X) + enable EAGLE for gfx950 in cookbook (#29194)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-25 03:54:48 -07:00
Mohammad Miadh Angkad e4976683f4 [Docs] Fix broken links in cookbook (#29261) 2026-06-25 02:31:55 -07:00