Xinyuan Tong
fa5c8a3101
[model] support encoder-free unified Text/Vision/Audio model ( #27167 )
2026-06-03 23:58:06 +08:00
e67810bea7
[SGLang Tracing] Add pd disaggregation mooncake backend tracing ( #23755 )
...
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com >
Co-authored-by: Shangming Cai <csmthu@gmail.com >
2026-06-03 16:43:29 +08:00
xutizhou
ac16dbf412
docs: add DeepSeek-V4 EPLB Waterfill tips ( #27049 )
2026-06-03 00:37:45 -07:00
Shaun Kotek
b8d7351a74
Feat/add w4a16 moe support to nemotron ( #25655 )
2026-06-02 22:42:26 -07:00
Mick
2d8cb87de7
docs: tag new diffusion cookbooks ( #27094 )
2026-06-03 09:50:27 +08:00
MingxuZh
b678448b8a
ci(xeon): merge 2 partitions into 1 job to reduce runner contention ( #26904 )
2026-06-03 09:46:27 +08:00
littleyellowbicycle
7271318dc3
【docs】The remote weight download function has been adjusted to be unsupported until the PTA interface is fixed. ( #27050 )
2026-06-02 20:55:23 +08:00
Mick
a6985e1be0
[diffusion] doc: add cookbook for lingbot-world ( #26958 )
2026-06-02 14:35:22 +08:00
Kurkur
f27fa0da93
[NPU][Docs] Kimi-K2.5 best practice ( #26774 )
2026-06-02 13:14:14 +08:00
1c0019da75
[Docs] GLM-4.7 cookbook: add NVIDIA Blackwell (B200, GB200) + NVFP4 sections ( #26384 )
...
Co-authored-by: Hao Phan <htphan@nvidia.com >
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-01 20:47:51 -07:00
Teng Ma and Zijie Xia
b562da0d9f
[PD] docs: clarify disaggregation IB device formats ( #25521 )
...
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai >
2026-06-02 11:38:33 +08:00
98a1b58c47
docs(cookbook): port popular model usage guides into cookbook pages ( #25813 )
...
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com >
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai >
2026-06-01 17:41:49 -07:00
Baizhou Zhang
2fdae94e46
docs: update RTX PRO 6000 deployment snippet ( #26968 )
2026-06-01 14:34:27 -07:00
eeecho and Claude Opus 4.6
524ba10eda
feat: SM120 (Blackwell Desktop) support for DeepSeek-V4 inference ( #24692 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-01 14:05:20 -07:00
a0670b5ba3
[SPEC] feat: add adaptive speculative decoding metrics ( #25940 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com >
2026-06-01 13:53:30 -07:00
Shu Wang and zijiexia
106092123f
Update Qwen3-Coder docs_new NVIDIA guidance ( #24435 )
...
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com >
2026-06-01 13:38:34 -07:00
Mick
1988a2c9ea
[diffusion] feat: improve cosmos3 serve API support ( #26926 )
2026-06-02 00:53:39 +08:00
Lukas Humbel and Claude Opus 4.7
d8a5a25c36
Refactor NIXL hicache. Add O_DIRECT support ( #25173 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-06-01 17:28:53 +02:00
amote-i
d078cb72bd
[NPU] [DOC] clarify Ascend NPU exclusive supported values for speculative args ( #26903 )
2026-06-01 16:40:11 +08:00
shadowxz109
4d20dc44fc
【NPU】add MiniMax2.5 best practice docs ( #26725 )
2026-06-01 10:09:12 +08:00
Haichuan Hu
acd689b407
feat: add SGLANG_RAY_BUNDLE_INDICES for fine-grained Ray bundle index control ( #24667 )
...
Signed-off-by: Haichuan Hu <kaisennhu@gmail.com >
2026-05-30 02:19:50 -07:00
silencejade
23a825c694
[DOC] [NPU] add qwen3.5-397b best practice to doc_new ( #26709 )
2026-05-30 14:20:24 +08:00
Chao Shi
6ce49e5f4c
[Utils] Support configure log level at runtime ( #26583 )
2026-05-29 14:49:06 -07:00
7fb7b41a3e
[docs] Qwen3.5 cookbook: multi-node, MTP TP overrides, dense mamba flag ( #26695 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-29 12:10:14 -07:00
Aditya Sharma and Xiaodong Ye
b2eed9e16d
[Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store ( #22868 )
...
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com >
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com >
2026-05-29 15:09:33 +08:00
Baizhou Zhang
69362cbc2c
[Doc] Update benchmark instruction for dsv4 ( #26668 )
2026-05-28 23:37:35 -07:00
40f91e6697
[Bugfix] [DSA] [Hisparse] Broadcast TP Rank 0 Topk Indexes to other TPs ( #24654 )
...
Co-authored-by: xz-keg <xuzou_keg@outlook.com >
Co-authored-by: xuzou <xu.zou@aminer.cn >
2026-05-28 21:14:46 -07:00
Yuhao Yang
a8cfae0b30
doc: update step-3.7-flash docker image tag ( #26625 )
2026-05-29 08:40:17 +08:00
3bdea78ad1
model: support Step-3.7-Flash ( #26565 )
...
Co-authored-by: yhyang201 <yhyang201@users.noreply.github.com >
Co-authored-by: luotingdan <luotingdan@stepfun.com >
2026-05-29 08:00:54 +08:00
Jimmy Shong
f838adb7d4
bench_serving: add Zipfian shared-prefix sampling to generated-shared-prefix ( #26378 )
2026-05-28 14:39:46 -07:00
97d129f8c6
# feat(bench): add SPEED-Bench dataset support to bench_serving ( #24149 )
...
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com >
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai >
2026-05-28 14:37:00 -07:00
Brayden Zhong and b8zhong
50e0b3b77f
Support Flashinfer Cute-DSL MLA attention ( #24737 )
...
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com >
2026-05-28 00:21:32 -07:00
Brayden Zhong
97dd6aad60
Add a little env var for disabling Flashinfer autotune cache ( #26193 )
2026-05-27 23:59:59 -07:00
Shaoting
14c1bb2721
[Feat][LMCache] Support LMCache mp mode ( #24089 )
...
Signed-off-by: Shaoting-Feng <stfeng@uw.edu >
2026-05-28 10:15:09 +08:00
Qiaolin Yu
561e54f803
Update kimi k25 launch command in cookbook ( #26511 )
2026-05-27 16:04:04 -07:00
loading66
a1ebc4917a
[NPU][DOCS]Add faq and feature Compatibilit ( #26464 )
2026-05-27 17:48:47 +08:00
Baizhou Zhang and Claude Opus 4.7
d6032c04b6
[docs] Fix V4 Pro balanced recipe ( #26451 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 23:50:26 -07:00
Zaili Wang
47617cc4df
[CPU Doc]Add Xeon CPU info in Qwen3 Cookbook ( #25971 )
2026-05-26 12:14:07 -07:00
zijiexia and Claude Opus 4.7
6afebc278a
[docs] DeepSeek-V4 cookbook: note cu129 image for GB200 Pro DeepEP backend ( #26413 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 12:08:48 -07:00
Baizhou Zhang
f0ba651d66
[Doc] Update pip install commands for Cuda12 ( #26344 )
2026-05-25 22:28:27 -07:00
Ziang Li
2b9dd9c8b3
[FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk ( #22851 )
2026-05-25 13:08:03 -07:00
Makcum888e and ronnie_zheng
0801cc05ed
[Diffusion][NPU] Disaggregation diffusion stages support for NPU ( #25895 )
...
Co-authored-by: ronnie_zheng <zl19940307@163.com >
2026-05-25 13:51:25 +03:00
Xiaoyu Zhang
533ef41112
[Diffusion] Default NVFP4 backend to FlashInfer TRTLLM ( #25523 )
2026-05-25 18:14:06 +08:00
Zhangheng and 晟海
a4db563c87
[hisparse]: update user guide ( #26249 )
...
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com >
2026-05-25 17:54:55 +08:00
Junlin Wu
aae04b1241
📝 docs(diffusion): add MXFP4 quantization docs ( #25904 )
2026-05-25 10:24:30 +03:00
zijiexia and Claude Opus 4.7
81cd338fcc
[docs] DeepSeek-V4 cookbook: balanced MegaMoE cap, H200 Pro FP4 mem-frac, nsa-* compat, PD-disagg fixes ( #26164 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-23 02:42:41 -07:00
longxin9715
c69844f043
[NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide ( #26069 )
2026-05-23 10:50:56 +08:00
jiayisunx
6339295556
[XPU] add apache-tvm-ffi dependency ( #26053 )
2026-05-22 16:09:08 +08:00
zijiexia and Claude Opus 4.7
88a37d7405
[docs] DeepSeek-V4 cookbook: split Quantization axis, add H100 SGLang FP8 ( #26057 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-22 00:52:10 -07:00
Alex O. P.
ae7c4226eb
[diffusion] model: support FLUX.2-klein-base ( #25661 )
2026-05-22 11:24:46 +08:00