Shaoting
|
14c1bb2721
|
[Feat][LMCache] Support LMCache mp mode (#24089)
Signed-off-by: Shaoting-Feng <stfeng@uw.edu>
|
2026-05-28 10:15:09 +08:00 |
|
Qiaolin Yu
|
561e54f803
|
Update kimi k25 launch command in cookbook (#26511)
|
2026-05-27 16:04:04 -07:00 |
|
loading66
|
a1ebc4917a
|
[NPU][DOCS]Add faq and feature Compatibilit (#26464)
|
2026-05-27 17:48:47 +08:00 |
|
 Baizhou ZhangandClaude Opus 4.7
|
d6032c04b6
|
[docs] Fix V4 Pro balanced recipe (#26451)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-26 23:50:26 -07:00 |
|
Zaili Wang
|
47617cc4df
|
[CPU Doc]Add Xeon CPU info in Qwen3 Cookbook (#25971)
|
2026-05-26 12:14:07 -07:00 |
|
 zijiexiaandClaude Opus 4.7
|
6afebc278a
|
[docs] DeepSeek-V4 cookbook: note cu129 image for GB200 Pro DeepEP backend (#26413)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-26 12:08:48 -07:00 |
|
Baizhou Zhang
|
f0ba651d66
|
[Doc] Update pip install commands for Cuda12 (#26344)
|
2026-05-25 22:28:27 -07:00 |
|
Ziang Li
|
2b9dd9c8b3
|
[FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk (#22851)
|
2026-05-25 13:08:03 -07:00 |
|
 Makcum888eandronnie_zheng
|
0801cc05ed
|
[Diffusion][NPU] Disaggregation diffusion stages support for NPU (#25895)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-25 13:51:25 +03:00 |
|
Xiaoyu Zhang
|
533ef41112
|
[Diffusion] Default NVFP4 backend to FlashInfer TRTLLM (#25523)
|
2026-05-25 18:14:06 +08:00 |
|
 Zhanghengand晟海
|
a4db563c87
|
[hisparse]: update user guide (#26249)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-05-25 17:54:55 +08:00 |
|
Junlin Wu
|
aae04b1241
|
📝 docs(diffusion): add MXFP4 quantization docs (#25904)
|
2026-05-25 10:24:30 +03:00 |
|
 zijiexiaandClaude Opus 4.7
|
81cd338fcc
|
[docs] DeepSeek-V4 cookbook: balanced MegaMoE cap, H200 Pro FP4 mem-frac, nsa-* compat, PD-disagg fixes (#26164)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-23 02:42:41 -07:00 |
|
longxin9715
|
c69844f043
|
[NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#26069)
|
2026-05-23 10:50:56 +08:00 |
|
jiayisunx
|
6339295556
|
[XPU] add apache-tvm-ffi dependency (#26053)
|
2026-05-22 16:09:08 +08:00 |
|
 zijiexiaandClaude Opus 4.7
|
88a37d7405
|
[docs] DeepSeek-V4 cookbook: split Quantization axis, add H100 SGLang FP8 (#26057)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-22 00:52:10 -07:00 |
|
Alex O. P.
|
ae7c4226eb
|
[diffusion] model: support FLUX.2-klein-base (#25661)
|
2026-05-22 11:24:46 +08:00 |
|
McZyWu
|
b2631a9a4d
|
[NPU] Docs op performance optimize (#25830)
|
2026-05-22 09:20:13 +08:00 |
|
 zijiexiaandClaude Opus 4.7
|
17dadebd4e
|
[Docs] DeepSeek-V4: switch H200 FP4 Pro to flashinfer_mxfp4, Flash Balanced too (#25923)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-21 13:51:40 -07:00 |
|
Yuhao Yang
|
81d686d9fa
|
Default MegaMoE to W4A8 for Max-Throughput recipe (#26004)
|
2026-05-21 11:54:16 -07:00 |
|
amote-i
|
ac83d8a339
|
docs: delete deprecated args from npu supported features (#25995)
|
2026-05-21 20:20:29 +08:00 |
|
loading66
|
2e0d2d4c18
|
[NPU][DOCS]Add best practice and benchmark result parameter description (#25875)
|
2026-05-21 19:08:10 +08:00 |
|
jianzhao-xu
|
f66881f03c
|
[NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#25384)
|
2026-05-21 17:32:25 +08:00 |
|
 jiayisunxandMa Mingfei
|
34479c19bd
|
[XPU] upgrade triton-xpu version to 3.7.1 (#25730)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-21 10:29:20 +08:00 |
|
Xiaoyu Zhang
|
ccbbae00ea
|
[codex] Reland Wan2.2 ModelOpt CI checkpoints (#25857)
|
2026-05-20 22:15:25 +08:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
8131641bc6
|
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-20 00:18:04 -07:00 |
|
Faradawn Yang
|
da6d549ab2
|
Update GLM-5 H200 FP8 (#25814)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
|
2026-05-20 14:44:54 +08:00 |
|
 Xinyuan TongandXinyuan Tong
|
52eebc82ae
|
[Docs] MiMo-V2.5 cookbook: B200 benchmarks + multi-layer EAGLE acceptance profile + long-context reference (#25359)
Co-authored-by: Xinyuan Tong <xinyuan.tong@radixark.ai>
|
2026-05-19 23:15:21 -07:00 |
|
Cheng Wan
|
a4b51d35ef
|
Revert "[codex] Update Wan2.2 ModelOpt CI checkpoints" (#25845)
|
2026-05-19 21:45:20 -07:00 |
|
Xiaoyu Zhang
|
80fc524809
|
[diffusion] quant: update Wan2.2 modelOpt CI checkpoints (#25483)
|
2026-05-20 09:05:39 +08:00 |
|
amote-i
|
de3fc46e3d
|
[NPU] [DOC] remove Qwen3-235B-A22B 2K+2K 100ms mixed mode benchmark (#25778)
|
2026-05-19 20:48:43 +08:00 |
|
Liangsheng Yin
|
e0273dcd31
|
pr-test-extra: re-trigger on labeled event (#25732)
|
2026-05-19 05:15:55 -07:00 |
|
 Arseniy MironovandNapkin-AI
|
45a85efc3a
|
[Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
|
2026-05-19 12:46:55 +03:00 |
|
amote-i
|
1f7bf155c3
|
[NPU] [DOCS] Improved the usability of Ascend NPU documents (#25735)
|
2026-05-19 16:22:22 +08:00 |
|
Ziang Li
|
78cb38ed5e
|
[FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918)
|
2026-05-19 01:04:48 -07:00 |
|
Kurkur
|
d028697d17
|
[NPU][Docs] Add Kimi-K2.5-W4A8 instance doc on NPU (#25269)
|
2026-05-19 09:08:28 +08:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
54eb2904a4
|
minor: docs include mac installation (#25178)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-05-18 15:48:59 +08:00 |
|
 Xia WeiwenandMa Mingfei
|
8d5ed330cc
|
[XPU] Enable qwen3.5 on XPU (#21668)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-18 14:59:19 +08:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
a080358cac
|
[Refactor] Refactor DeepEP dispatcher (#22822)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-05-18 04:36:42 +03:00 |
|
Baizhou Zhang
|
6dcacb1159
|
[Doc] Fix several places for dpsk v4 cookbook (#25506)
|
2026-05-16 21:54:15 -07:00 |
|
Yuhao Yang
|
57eb5bdaf6
|
[Doc] DSV4 cookbook: clean up env vars, add MegaMoE toggle, unify docker image (#25412)
|
2026-05-16 11:28:05 -07:00 |
|
zijiexia
|
9f26697d6a
|
[Docs] Update DeepSeek V4 cookbook to use the latest docker image (#25410)
|
2026-05-16 11:17:51 -07:00 |
|
Zheng Luo
|
435ea41cf0
|
Delegate ModelExpress loading to package (#24723)
Signed-off-by: Zheng Luo <zheluo@nvidia.com>
|
2026-05-16 11:16:44 -07:00 |
|
Xinyuan Tong
|
33f1d3915f
|
[NEW MODEL] Add H200 validation for Ring-2.6-1T cookbook (#25370)
|
2026-05-15 11:47:15 -07:00 |
|
Baizhou Zhang
|
1a1d69507d
|
[Doc] Update MegaMoE usage (#25378)
|
2026-05-15 02:17:50 -07:00 |
|
Zhangheng
|
c7e879e43f
|
Add hicache feature in dsv4 cookbook (#25369)
|
2026-05-15 00:06:44 -07:00 |
|
Xinyuan Tong
|
c3daa77e9a
|
[NEW MODEL] Add Ring-2.6-1T cookbook (#25360)
|
2026-05-14 23:32:22 -07:00 |
|
 Haoguang CaiandClaude Sonnet 4.6
|
626fd61308
|
📝 docs: add canonical URL to fix Google indexing lmsysorg.mintlify.app instead of docs.sglang.io (#24935)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-14 16:46:40 -07:00 |
|
Liangsheng Yin
|
67096f48bf
|
Revert "[MoE] Decouple Mega MoE from DeepEP backend" (#25317)
|
2026-05-14 16:00:41 -07:00 |
|
Yuhao Yang
|
37f030a0de
|
[MoE] Decouple Mega MoE from DeepEP backend (#24884)
|
2026-05-15 02:01:44 +08:00 |
|