Commit Graph
120 Commits
Author SHA1 Message Date
Liangsheng Yin c0480a88be [Spec] Retire Spec V1 (#27964) 2026-06-11 16:15:15 -07:00
Trevor Morris 0bac184425 [NVIDIA] Update Minimax-M2.5,M2.7 docs with flags for performance (#24465) 2026-06-11 14:58:44 -07:00
Brian Chao 7f57b344c9 [diffusion] feat: progressive resolution growing for Ideogram 4 via GPU DCT upsampling with up to 1.56× speedup (#27736) 2026-06-11 23:16:53 +08:00
Chi McIsaacandMick b2728bda9d [diffusion] feat: use fused w8a8 kernel for Ideogram4 weight-only linear as an opt-in (#27590)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-11 23:15:27 +08:00
ming_wang ce0ff154a5 add mimo best practice (#27665) 2026-06-11 10:15:31 +08:00
amote-i fc1fee528c [NPU] [DOC] replace <code> with backticks and remove obsolete params (#27677) 2026-06-11 09:33:31 +08:00
Baizhou Zhang 3c1b0fb226 [1/n] [CP] Simplify prefill context parallel server args (#27312) 2026-06-10 14:11:38 -07:00
billishyahaoandHAI 0ae27405d0 [AMD] Support eplb for moriep (#22985)
Co-authored-by: HAI <hixiao@gmail.com>
2026-06-10 10:23:51 -07:00
1cf8efdd08 docs: Diffusion Gemma cookbook (#27824)
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 09:33:01 -07:00
Ziang Li 01f10acd06 Implement online nvfp4 quantization (#26083) 2026-06-10 00:26:51 -07:00
Mick e8a437ef26 [diffusion] doc: update docs architecture (#27767) 2026-06-10 14:18:10 +08:00
2495c02c2c [Refactor] Cuda Graph Runner/Backend Refactor (#23906)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-09 21:36:57 -07:00
bcd9c5a903 update pytorch-xpu to 2.12 (#27133)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
2026-06-10 09:25:30 +08:00
Liangsheng Yin 186f1e300a [CI] Move JIT kernel tests + benchmarks to test/registered/jit; add in-package guard (#27644) 2026-06-09 12:37:39 -07:00
Brian Chao b0cd533a96 [diffusion] feat: progressive resolution growing for image and video models (#27524) 2026-06-08 20:44:07 +08:00
52a5c01eba [plugin] enable OOT platforms to provide custom quant configs (#25347)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <laldevashish@gmail.com>
2026-06-07 12:48:55 +08:00
Артем Савкин 9f28512737 [NPU] Update documentation for software version upgrades (#26731) 2026-06-06 15:06:13 +03:00
EduardDurech e9dbbd19e9 [model] Apertus Tool/Function and Reasoning parser (#25100) 2026-06-06 00:04:31 -07:00
6b180959a8 [SPEC][5/N] feat: batchsize-aware support for adaptive speculative_num_steps (#24055)
Co-authored-by: 坤钧 <maoyuhan.myh@antgroup.co>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: shuwenn <2508695655@qq.com>
2026-06-05 15:43:02 -07:00
McZyWu d8487bad06 Update best practice for qwen3-next-80b-a3b-instruct (#27353) 2026-06-05 17:02:38 +08:00
jianzhao-xu 4248695b07 [NPU] add GLM model best practice docs (#27032) 2026-06-05 14:27:19 +08:00
zijiexiaandClaude Opus 4.8 c6c1f1a29a docs: sync legacy docs/-only updates into docs_new (Mintlify) (#27308)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 19:45:13 -07:00
e76d36214b Changes for SM120 perf and usability for NVFP4 (#26496)
Co-authored-by: Martin Vit <martin@voipmonitor.org>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
2026-06-04 15:29:25 -07:00
Ziang Li 4cfebbb95f [FlashInfer v0.6.12] Support FlashInfer 4over6 NVFP4 (#25239) 2026-06-04 14:35:07 -07:00
Jinyan ChenandJinyan Chen 10b6b45cad docs: add DeepSeek V4 FP4 indexer usage (#27035)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
2026-06-04 00:44:21 -07:00
293816ab14 [AMD][MXFP4] Online MXFP4 quantization 1/N - dense and MOE models w. original BF16 weight (#18005)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: Colin Zeng <Colin.Zeng@amd.com>
2026-06-03 12:55:24 -07:00
e67810bea7 [SGLang Tracing] Add pd disaggregation mooncake backend tracing (#23755)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-03 16:43:29 +08:00
Shaun Kotek b8d7351a74 Feat/add w4a16 moe support to nemotron (#25655) 2026-06-02 22:42:26 -07:00
MingxuZh b678448b8a ci(xeon): merge 2 partitions into 1 job to reduce runner contention (#26904) 2026-06-03 09:46:27 +08:00
littleyellowbicycle 7271318dc3 【docs】The remote weight download function has been adjusted to be unsupported until the PTA interface is fixed. (#27050) 2026-06-02 20:55:23 +08:00
Kurkur f27fa0da93 [NPU][Docs] Kimi-K2.5 best practice (#26774) 2026-06-02 13:14:14 +08:00
Teng MaandZijie Xia b562da0d9f [PD] docs: clarify disaggregation IB device formats (#25521)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-02 11:38:33 +08:00
98a1b58c47 docs(cookbook): port popular model usage guides into cookbook pages (#25813)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-06-01 17:41:49 -07:00
a0670b5ba3 [SPEC] feat: add adaptive speculative decoding metrics (#25940)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com>
2026-06-01 13:53:30 -07:00
Lukas HumbelandClaude Opus 4.7 d8a5a25c36 Refactor NIXL hicache. Add O_DIRECT support (#25173)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 17:28:53 +02:00
amote-i d078cb72bd [NPU] [DOC] clarify Ascend NPU exclusive supported values for speculative args (#26903) 2026-06-01 16:40:11 +08:00
shadowxz109 4d20dc44fc 【NPU】add MiniMax2.5 best practice docs (#26725) 2026-06-01 10:09:12 +08:00
Haichuan Hu acd689b407 feat: add SGLANG_RAY_BUNDLE_INDICES for fine-grained Ray bundle index control (#24667)
Signed-off-by: Haichuan Hu <kaisennhu@gmail.com>
2026-05-30 02:19:50 -07:00
silencejade 23a825c694 [DOC] [NPU] add qwen3.5-397b best practice to doc_new (#26709) 2026-05-30 14:20:24 +08:00
Chao Shi 6ce49e5f4c [Utils] Support configure log level at runtime (#26583) 2026-05-29 14:49:06 -07:00
Aditya SharmaandXiaodong Ye b2eed9e16d [Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-05-29 15:09:33 +08:00
40f91e6697 [Bugfix] [DSA] [Hisparse] Broadcast TP Rank 0 Topk Indexes to other TPs (#24654)
Co-authored-by: xz-keg <xuzou_keg@outlook.com>
Co-authored-by: xuzou <xu.zou@aminer.cn>
2026-05-28 21:14:46 -07:00
Jimmy Shong f838adb7d4 bench_serving: add Zipfian shared-prefix sampling to generated-shared-prefix (#26378) 2026-05-28 14:39:46 -07:00
97d129f8c6 # feat(bench): add SPEED-Bench dataset support to bench_serving (#24149)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-05-28 14:37:00 -07:00
Brayden Zhongandb8zhong 50e0b3b77f Support Flashinfer Cute-DSL MLA attention (#24737)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-28 00:21:32 -07:00
Brayden Zhong 97dd6aad60 Add a little env var for disabling Flashinfer autotune cache (#26193) 2026-05-27 23:59:59 -07:00
Shaoting 14c1bb2721 [Feat][LMCache] Support LMCache mp mode (#24089)
Signed-off-by: Shaoting-Feng <stfeng@uw.edu>
2026-05-28 10:15:09 +08:00
loading66 a1ebc4917a [NPU][DOCS]Add faq and feature Compatibilit (#26464) 2026-05-27 17:48:47 +08:00
Baizhou Zhang f0ba651d66 [Doc] Update pip install commands for Cuda12 (#26344) 2026-05-25 22:28:27 -07:00
Ziang Li 2b9dd9c8b3 [FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk (#22851) 2026-05-25 13:08:03 -07:00