Commit Graph
188 Commits
Author SHA1 Message Date
3a679459e5 [bench] Add agentic-trace multi-turn dataset to bench_serving (#29215)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 19:45:44 -07:00
Junlin Wuandronnie_zheng 3abdbab9bb ✨ [llm][npu][quant] Add W4A8 MXFP quantization support for Qwen3 Dense on Ascend NPU (#23650)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-06 19:23:26 +03:00
Baizhou Zhang 5eb1b6a7ba Remove retired DSA env paths (#29912) 2026-07-05 22:58:02 -07:00
addffd7489 [Diffusion] Diffusion model support log-requests (#23049)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-05 12:01:26 +03:00
Cao E 9df16b5ba9 [XPU] Remove redundant xpu graph backend and make xpu graph opt-in by default (#29911) 2026-07-03 15:58:56 +08:00
amote-i 2fe7182e75 [DOC] [NPU] update supported features on ascend npu (#30011) 2026-07-03 15:36:41 +08:00
ming_wang d8a4f7a7aa add mimo-v2-flash model tutorial (#29932) 2026-07-03 11:02:24 +08:00
Peng Xingchen a6bc7fef90 glm5.2 on ascend doc (new version) (#29828) 2026-07-03 10:19:44 +08:00
amote-i 70a813493f [NPU] [DOC] add missing DEEP_NORMAL_MODE_USE_INT8_QUANT for w8a8+deepep scenarios (#29937) 2026-07-03 10:18:22 +08:00
Cao E 926140d789 [XPU] Enable XPU graph support (decode full-graph + prefill tc_piecewise) (#29053) 2026-07-02 13:24:35 +08:00
Zaili Wang cb06c4e6ce [CPU] Fix model failures on Xeon (#29497) 2026-07-02 13:20:18 +08:00
Cheng Wanandlch1475369 4a8e76805c feat(mem_cache): unified memory pool for hybrid Mamba / SWA models (#29678)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-07-01 13:21:59 -07:00
qinsir5522 a7390b17f8 [NPU]Modify --lora-backend & --moe-runner-backend description. (#29793) 2026-07-01 11:15:46 +08:00
97fc4dfd73 [Doc]Checking and modifying Markdown formatting issues and link validity (#28586)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-30 20:27:38 +03:00
Trevor Morris b8c25bfaa7 [NVIDIA] Support flashinfer a2a with flashinfer_trtllm_routed moe (#22394) 2026-06-29 16:23:58 -07:00
Cheng Wanandlch1475369 fc96edd297 feat(mem_cache): page-major (layer-major within a page) KV/state layout (#29533)
Co-authored-by: lch1475369 <lch1475369@gmail.com>
2026-06-29 14:49:54 -07:00
zijiexiaandClaude Opus 4.8 11b7ed7c9e [Docs] Add --prerelease=allow so uv installs the latest sglang (#29676)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 14:13:34 -07:00
Jyothirmai KottuandXinyuan Tong 473a278dd1 model: support nvidia/LocateAnything-3B (#28958)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-30 00:16:42 +08:00
amote-i 489017b3d6 [NPU] [DOC] Update deterministic inference feature support status to A2, A3 (#29632) 2026-06-29 19:17:06 +08:00
danielafrimiandDaniel Afrimi a2b5ce2ed1 Add stochastic rounding for FP16 Mamba SSM cache (#26929)
Signed-off-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
Co-authored-by: Daniel Afrimi <dafrimi@login-lyris01.lyris.clusters.nvidia.com>
2026-06-29 01:47:09 -07:00
jianzhao-xu 2260e612f6 [NPU] update best practicce docs from testcase (#29492) 2026-06-29 11:30:27 +08:00
Liangsheng Yin 909123ddb8 [misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591) 2026-06-28 18:38:28 -07:00
b030b1a5f3 hisparse: support NIXL DRAM KV destinations for HiSparse (#27563)
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-27 22:32:29 +08:00
jonah-bermanandXinyuan Tong 2f34dbe372 Add native Exa-backed web_search support (#29342)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-27 14:41:09 +01:00
amote-i a3c5e286f6 [NPU] [DOC] Fix and update Ascend NPU docs (#29501) 2026-06-27 18:23:11 +08:00
Yaochen Hanandronnie_zheng c98d31143d update quantization code owner and document quantization contributions (#26784)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-26 19:40:33 +03:00
jianzhao-xu cc294829aa [NPU] fix best practicce docs (#29303) 2026-06-26 11:17:50 +08:00
cfc0a0e0e0 Add Intel Quantization Support in SGLang (#18139)
Signed-off-by: Mengni Wang <mengni.wang@intel.com>
Signed-off-by: WeiweiZhang1 <weiwei1.zhang@intel.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Weiwei <weiwei1.zhang@intel.com>
2026-06-26 09:54:35 +08:00
amote-i 10ff3c1dcb [NPU] [DOC] Add environment prerequisites to model tutorials (#29293) 2026-06-26 09:51:06 +08:00
Yi Zhong f83cbc2516 Add LFM2.5-230M to the LFM2.5 cookbook (#29321) 2026-06-26 08:27:57 +08:00
52c32035eb [diffusion] Add Qwen-Image ModelOpt NVFP4 support (#28928)
Co-authored-by: jingyu-ml <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 22:56:09 +08:00
Mick 890b38c211 [diffusion] doc: fix diffusion docs and cookbook drift (#29302) 2026-06-25 21:13:00 +08:00
Brayden ZhongandBrayden Zhong f82addd4a8 Support online MXFP8 quantization for ungated MoE (#27939)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-24 16:58:48 -07:00
amote-i 73d976e375 [NPU] [DOC] Fix TOC of Ascend NPU Docs (#29129) 2026-06-24 16:47:10 +08:00
zijiexiaandClaude Opus 4.8 dd2d919e21 [Docs] Fix mem-fraction-static default and document how it is computed (#29135)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 08:05:06 +00:00
Trevor Morris f74a1722e6 [NVIDIA] Support TF32 matmul to improve MiniMax gate gemm performance (#22744) 2026-06-23 14:54:54 -07:00
amote-i 84338df6f0 [NPU] [DOC] Update contribution guide of Ascend NPU (#28909) 2026-06-23 10:00:12 +08:00
99c18cceec Sync server arguments and environment variables + update various documentation (#28674)
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-06-22 10:23:46 -07:00
jianzhao-xu 4e1d25117b [NPU] update best practice docs from testcase (#28621) 2026-06-22 16:58:48 +08:00
amote-i 93553a67a3 [NPU] [DOC] Create deployment tutorials for mainstream models on Ascend NPU (#27893) 2026-06-22 16:21:13 +08:00
018d0c21dc [Docs] Add Anthropic-compatible API documentation (#28522)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 04:01:09 +00:00
amote-i 5deca2d39f [DOC] [NPU] Update features on Ascend NPU (#28643) 2026-06-22 09:50:51 +08:00
cctry 6d4ca9bc54 Cap SWA pool sizing with chunk cache (#28755) 2026-06-21 01:06:59 -07:00
Oguz Ulgen 3af991fb3e [AMD] Make breakable CUDA graph run on ROCm/HIP (#28173) 2026-06-19 07:16:00 -07:00
Clintandclintg6 fac11f3bc1 [AMD] Document Mori XGMI for Single-Node PD Disaggregation (#25094)
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
2026-06-18 19:25:49 -07:00
Mick 05b3fd0f44 [diffusion] chore: remove ltx2 snapshot mode (#28533) 2026-06-18 10:20:21 +08:00
zijiexia 74e2e48c82 Introduce CpuDeviceMixin and CpuSRTPlatform (#26385) 2026-06-17 17:41:17 -07:00
chenxu214 71b090a8e7 [Ascend]GLM 5.2 deployment (#28433) 2026-06-17 09:55:50 +08:00
amote-i 74155d32bd [NPU] [DOC] Update Ascend NPU docs: HDK 25.5.2, Triton 3.2.1.dev20260530 (#28311) 2026-06-17 09:38:31 +08:00
Jyothirmai Kottu b8b8992dde docs: add Amazon SageMaker AI deployment guide (#28338) 2026-06-16 13:33:22 -07:00