Commit Graph
102 Commits
Author SHA1 Message Date
6b180959a8 [SPEC][5/N] feat: batchsize-aware support for adaptive speculative_num_steps (#24055)
Co-authored-by: 坤钧 <maoyuhan.myh@antgroup.co>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
2026-06-05 15:43:02 -07:00
McZyWu d8487bad06 Update best practice for qwen3-next-80b-a3b-instruct (#27353) 2026-06-05 17:02:38 +08:00
jianzhao-xu 4248695b07 [NPU] add GLM model best practice docs (#27032) 2026-06-05 14:27:19 +08:00
zijiexiaandClaude Opus 4.8 c6c1f1a29a docs: sync legacy docs/-only updates into docs_new (Mintlify) (#27308)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 19:45:13 -07:00
e76d36214b Changes for SM120 perf and usability for NVFP4 (#26496)
Co-authored-by: Martin Vit <martin@voipmonitor.org>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-04 15:29:25 -07:00
Ziang Li 4cfebbb95f [FlashInfer v0.6.12] Support FlashInfer 4over6 NVFP4 (#25239) 2026-06-04 14:35:07 -07:00
Jinyan ChenandJinyan Chen 10b6b45cad docs: add DeepSeek V4 FP4 indexer usage (#27035)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
2026-06-04 00:44:21 -07:00
293816ab14 [AMD][MXFP4] Online MXFP4 quantization 1/N - dense and MOE models w. original BF16 weight (#18005)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: Colin Zeng <Colin.Zeng@amd.com>
2026-06-03 12:55:24 -07:00
e67810bea7 [SGLang Tracing] Add pd disaggregation mooncake backend tracing (#23755)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-03 16:43:29 +08:00
Shaun Kotek b8d7351a74 Feat/add w4a16 moe support to nemotron (#25655) 2026-06-02 22:42:26 -07:00
MingxuZh b678448b8a ci(xeon): merge 2 partitions into 1 job to reduce runner contention (#26904) 2026-06-03 09:46:27 +08:00
littleyellowbicycle 7271318dc3 【docs】The remote weight download function has been adjusted to be unsupported until the PTA interface is fixed. (#27050) 2026-06-02 20:55:23 +08:00
Kurkur f27fa0da93 [NPU][Docs] Kimi-K2.5 best practice (#26774) 2026-06-02 13:14:14 +08:00
Teng MaandZijie Xia b562da0d9f [PD] docs: clarify disaggregation IB device formats (#25521)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-02 11:38:33 +08:00
98a1b58c47 docs(cookbook): port popular model usage guides into cookbook pages (#25813)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-01 17:41:49 -07:00
a0670b5ba3 [SPEC] feat: add adaptive speculative decoding metrics (#25940)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com>
2026-06-01 13:53:30 -07:00
Lukas HumbelandClaude Opus 4.7 d8a5a25c36 Refactor NIXL hicache. Add O_DIRECT support (#25173)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 17:28:53 +02:00
amote-i d078cb72bd [NPU] [DOC] clarify Ascend NPU exclusive supported values for speculative args (#26903) 2026-06-01 16:40:11 +08:00
shadowxz109 4d20dc44fc 【NPU】add MiniMax2.5 best practice docs (#26725) 2026-06-01 10:09:12 +08:00
Haichuan Hu acd689b407 feat: add SGLANG_RAY_BUNDLE_INDICES for fine-grained Ray bundle index control (#24667)
Signed-off-by: Haichuan Hu <kaisennhu@gmail.com>
2026-05-30 02:19:50 -07:00
silencejade 23a825c694 [DOC] [NPU] add qwen3.5-397b best practice to doc_new (#26709) 2026-05-30 14:20:24 +08:00
Chao Shi 6ce49e5f4c [Utils] Support configure log level at runtime (#26583) 2026-05-29 14:49:06 -07:00
Aditya SharmaandXiaodong Ye b2eed9e16d [Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-05-29 15:09:33 +08:00
40f91e6697 [Bugfix] [DSA] [Hisparse] Broadcast TP Rank 0 Topk Indexes to other TPs (#24654)
Co-authored-by: xz-keg <xuzou_keg@outlook.com>
Co-authored-by: xuzou <xu.zou@aminer.cn>
2026-05-28 21:14:46 -07:00
Jimmy Shong f838adb7d4 bench_serving: add Zipfian shared-prefix sampling to generated-shared-prefix (#26378) 2026-05-28 14:39:46 -07:00
97d129f8c6 # feat(bench): add SPEED-Bench dataset support to bench_serving (#24149)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-05-28 14:37:00 -07:00
Brayden Zhongandb8zhong 50e0b3b77f Support Flashinfer Cute-DSL MLA attention (#24737)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-28 00:21:32 -07:00
Brayden Zhong 97dd6aad60 Add a little env var for disabling Flashinfer autotune cache (#26193) 2026-05-27 23:59:59 -07:00
Shaoting 14c1bb2721 [Feat][LMCache] Support LMCache mp mode (#24089)
Signed-off-by: Shaoting-Feng <stfeng@uw.edu>
2026-05-28 10:15:09 +08:00
loading66 a1ebc4917a [NPU][DOCS]Add faq and feature Compatibilit (#26464) 2026-05-27 17:48:47 +08:00
Baizhou Zhang f0ba651d66 [Doc] Update pip install commands for Cuda12 (#26344) 2026-05-25 22:28:27 -07:00
Ziang Li 2b9dd9c8b3 [FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk (#22851) 2026-05-25 13:08:03 -07:00
Makcum888eandronnie_zheng 0801cc05ed [Diffusion][NPU] Disaggregation diffusion stages support for NPU (#25895)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-25 13:51:25 +03:00
Xiaoyu Zhang 533ef41112 [Diffusion] Default NVFP4 backend to FlashInfer TRTLLM (#25523) 2026-05-25 18:14:06 +08:00
Zhanghengand晟海 a4db563c87 [hisparse]: update user guide (#26249)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-05-25 17:54:55 +08:00
Junlin Wu aae04b1241 📝 docs(diffusion): add MXFP4 quantization docs (#25904) 2026-05-25 10:24:30 +03:00
longxin9715 c69844f043 [NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#26069) 2026-05-23 10:50:56 +08:00
jiayisunx 6339295556 [XPU] add apache-tvm-ffi dependency (#26053) 2026-05-22 16:09:08 +08:00
Alex O. P. ae7c4226eb [diffusion] model: support FLUX.2-klein-base (#25661) 2026-05-22 11:24:46 +08:00
McZyWu b2631a9a4d [NPU] Docs op performance optimize (#25830) 2026-05-22 09:20:13 +08:00
amote-i ac83d8a339 docs: delete deprecated args from npu supported features (#25995) 2026-05-21 20:20:29 +08:00
loading66 2e0d2d4c18 [NPU][DOCS]Add best practice and benchmark result parameter description (#25875) 2026-05-21 19:08:10 +08:00
jianzhao-xu f66881f03c [NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#25384) 2026-05-21 17:32:25 +08:00
jiayisunxandMa Mingfei 34479c19bd [XPU] upgrade triton-xpu version to 3.7.1 (#25730)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-21 10:29:20 +08:00
Xiaoyu Zhang ccbbae00ea [codex] Reland Wan2.2 ModelOpt CI checkpoints (#25857) 2026-05-20 22:15:25 +08:00
Cheng WanandClaude Sonnet 4.6 8131641bc6 [Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-05-20 00:18:04 -07:00
Cheng Wan a4b51d35ef Revert "[codex] Update Wan2.2 ModelOpt CI checkpoints" (#25845) 2026-05-19 21:45:20 -07:00
Xiaoyu Zhang 80fc524809 [diffusion] quant: update Wan2.2 modelOpt CI checkpoints (#25483) 2026-05-20 09:05:39 +08:00
amote-i de3fc46e3d [NPU] [DOC] remove Qwen3-235B-A22B 2K+2K 100ms mixed mode benchmark (#25778) 2026-05-19 20:48:43 +08:00
Liangsheng Yin e0273dcd31 pr-test-extra: re-trigger on labeled event (#25732) 2026-05-19 05:15:55 -07:00